Publication:

Semi-supervised and Representation Learning for Improved Classification and Stratification in EHR Data

Loading...
Thumbnail Image

Date

2025-05-13

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Wang, Linshanshan. 2025. Semi-supervised and Representation Learning for Improved Classification and Stratification in EHR Data. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

The rapid digitization of healthcare has given rise to vast repositories of electronic health record (EHR) data, offering unprecedented opportunities for data-driven advancements in disease prediction, patient stratification, and clinical decision-making. However, the high dimensionality, sparsity, and heterogeneity of EHR data present unique statistical and computational challenges. Moreover, the scarcity of high-quality labels—due to the cost and complexity of manual annotation—further complicates supervised modeling efforts. This dissertation addresses these challenges through a unified framework of semi-supervised learning and representation learning for improved classification and stratification in EHR data, with applications to phenotyping, disability prediction, and patient subgroup discovery.

The overarching goal of this work is to develop scalable, robust, and interpretable methods that leverage both labeled and unlabeled EHR data, improve generalizability across populations, and uncover clinically meaningful structure in complex disease settings. The dissertation is composed of three interrelated papers, each tackling a key methodological bottleneck in modern EHR-based machine learning: (1) evaluating model performance under distributional shift, (2) learning rich patient representations in the presence of limited labels, and (3) stratifying heterogeneous patient populations using outcome-informed embeddings.

In Chapter 1, we consider the problem of evaluating the performance of binary classifiers when labeled data are unavailable in a target population. This setting is common in clinical phenotyping tasks, where models are trained using limited chart-reviewed labels in one cohort and then applied to other cohorts with potentially different covariate distributions. We propose STEAM Semi-supervised Transfer lEarning of Accuracy Measures), a doubly robust estimation procedure for receiver operating characteristic (ROC) parameters under covariate shift. STEAM combines calibrated density ratio weighting with robust outcome imputation, using both unlabeled source and target data to improve efficiency while protecting against model misspecification. Through theoretical guarantees and empirical results, we demonstrate that STEAM enables accurate performance assessment in unlabeled target populations, with applications to phenotyping models in rheumatoid arthritis on temporally evolving EHR cohort.

Building on the challenge of label scarcity, Chapter 2 shifts focus to semi-supervised representation learning for predictive modeling. We propose SCORE (Semi-supervised Clustering thrOugh REp- resentation learning), a generative embedding framework that models the joint distribution of high-dimensional EHR features using a multivariate Poisson-LogNormal distribution, with pretrained code embeddings capturing semantic relationships between clinical concepts. SCORE integrates limited labeled data via a hybrid Expectation-Maximization and Gaussian Variational Approximation algorithm, enabling efficient and theoretically sound inference in large-scale, partially labeled cohorts. We show that SCORE produces informative and transferable patient embeddings, improving prediction of disability status in multiple sclerosis (MS) and outperforming conventional supervised and unsupervised methods.

Finally, Chapter 3 addresses the critical task of patient stratification in heterogeneous diseases. We focus on Alzheimer’s disease (AD), where progression and prognosis vary substantially with age. We propose SOLAR (age-Specific Outcome-guided representation Learning for pAtient clusteRing), a novel clustering framework that incorporates time-to-event outcomes and explicitly models age-group structure using a multitask learning paradigm. SOLAR jointly learns low-dimensional patient representations across age groups, encouraging shared structure while allowing age-specific flexibility. By integrating survival information and modeling age-related heterogeneity, SOLAR identifies clinically meaningful AD subtypes with distinct prognostic profiles, improving both interpretability and clinical utility over existing age-unaware or outcome-agnostic methods.

Together, these three works present a cohesive framework for semi-supervised and representation learning in EHR analysis. The methods developed here contribute new strategies for evaluating, predicting, and stratifying patient outcomes in data-scarce, high-dimensional clinical settings. In doing so, they aim to advance the broader goals of personalized medicine and evidence-based healthcare by making machine learning more robust, scalable, and clinically relevant.

Description

Other Available Sources

Research Data

Keywords

Classification, EHR data, Representation Learning, Semi-supervised Learning, Biostatistics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories