Publication:

A Comparison of Natural Language Models to Subtype Ischemic Stroke from Electronic Health Records

Loading...
Thumbnail Image

Date

2025-03-17

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Alemu, Kidist Workeye. 2023. A Comparison of Natural Language Models to Subtype Ischemic Stroke from Electronic Health Records. Bachelors Thesis, Harvard University Engineering and Applied Sciences.

Abstract

Ischemic stroke is a leading cause of disability and death worldwide. While the prevalence of ischemic stroke varies across race and ethnicity, it is particularly pronounced in low-resource medical institutions, where patients may experience higher rates of post-stroke complications and mortality. Despite concerted attempts to implement established interventions and explore novel treatment modalities, the coarse classification of ischemic stroke obscures the underlying heterogeneity in both pathophysiological mechanisms and clinical manifestations, rendering these efforts insufficient. As such, there is huge value in investigating the underlying etiology of ischemic stroke in large, diverse cohorts of patients that could power refined subtype discovery. While electronic health record (EHR) data represent a valuable resource for stroke subtyping, there are few studies that have utilized EHR data for building these cohorts due to the significant human effort required to label features and cases.

To this goal, the rapid progress of Natural Language Processing methods has opened the door to fully automated stroke subtyping. This study investigates the performance of two Natural Language Processing approaches, namely Logistic Regression, a statistical model, and Clinical Longformer, a pre-trained transformer model, in subtyping ischemic stroke directly from the EHR. The models were trained and tested on EHR data of about 3000 stroke patients adjudicated by board-certified neurologists. Both models achieve commendable performance, with the transformer model slightly outperforming the statistical model with recall and precision of 0.83 and 0.74 respectively. This finding highlights that Natural Language Processing offers a more consistent and scalable approach to subtyping ischemic stroke from EHR, which could significantly enhance the statistical power and facilitate large-scale stroke research. In particular, this outcome has enabled us to infer toast subtypes across a sizable cohort of 30,000 coded strokes at Massachusetts General Hospital, thereby opening up novel avenues for investigating stroke risk prediction and genetic underpinnings.

Description

Other Available Sources

Research Data

Keywords

Computer science, Neurosciences

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories