Batra, PuneetAlemu, Kidist Workeye2025-09-1820232025-03-172023Alemu, Kidist Workeye. 2023. A Comparison of Natural Language Models to Subtype Ischemic Stroke from Electronic Health Records. Bachelors Thesis, Harvard University Engineering and Applied Sciences.30315509https://dash.harvard.edu/handle/1/42719551Ischemic stroke is a leading cause of disability and death worldwide. While the prevalence of ischemic stroke varies across race and ethnicity, it is particularly pronounced in low-resource medical institutions, where patients may experience higher rates of post-stroke complications and mortality. Despite concerted attempts to implement established interventions and explore novel treatment modalities, the coarse classification of ischemic stroke obscures the underlying heterogeneity in both pathophysiological mechanisms and clinical manifestations, rendering these efforts insufficient. As such, there is huge value in investigating the underlying etiology of ischemic stroke in large, diverse cohorts of patients that could power refined subtype discovery. While electronic health record (EHR) data represent a valuable resource for stroke subtyping, there are few studies that have utilized EHR data for building these cohorts due to the significant human effort required to label features and cases. To this goal, the rapid progress of Natural Language Processing methods has opened the door to fully automated stroke subtyping. This study investigates the performance of two Natural Language Processing approaches, namely Logistic Regression, a statistical model, and Clinical Longformer, a pre-trained transformer model, in subtyping ischemic stroke directly from the EHR. The models were trained and tested on EHR data of about 3000 stroke patients adjudicated by board-certified neurologists. Both models achieve commendable performance, with the transformer model slightly outperforming the statistical model with recall and precision of 0.83 and 0.74 respectively. This finding highlights that Natural Language Processing offers a more consistent and scalable approach to subtyping ischemic stroke from EHR, which could significantly enhance the statistical power and facilitate large-scale stroke research. In particular, this outcome has enabled us to infer toast subtypes across a sizable cohort of 30,000 coded strokes at Massachusetts General Hospital, thereby opening up novel avenues for investigating stroke risk prediction and genetic underpinnings.application/pdfenComputer scienceNeurosciencesA Comparison of Natural Language Models to Subtype Ischemic Stroke from Electronic Health RecordsThesis or Dissertation2025-09-18