Publication:

Statistical Methods for Negative-Unlabeled Data with Application to Long COVID

Loading...
Thumbnail Image

Date

2025-09-06

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Cao, Tingyi. 2025. Statistical Methods for Negative-Unlabeled Data with Application to Long COVID. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Negative-unlabeled data arise in settings where a subset of observations has known negative outcomes (e.g., people without a history of SARS-CoV-2 infection do not have Long COVID), while the remaining observations are unlabeled (e.g., individuals with prior infection have uncertain Long COVID status). This partial labeling structure presents challenges for statistical inference, particularly in characterizing heterogeneous conditions such as Long COVID (LC). Existing supervised, unsupervised, or semi-supervised approaches cannot directly accommodate negative-unlabeled data. However, data with this structure are increasingly common in public health, including electronic health records and post-infectious syndrome research. In this dissertation, we develop statistical approaches tailored for negative-unlabeled data with applications to LC and potential extensions to other partially labeled disease phenotyping problems.

LC is a multisystem condition with variable pathophysiological manifestations. Its fluctuating symptom patterns can persist for months or years following acute SARS-CoV-2 infection. Understanding its clinical heterogeneity and longitudinal trajectories remains an urgent research priority. This work is motivated by the NIH-sponsored Researching COVID to Enhance Recovery (RECOVER) Initiative, a large observational cohort study of individuals followed quarterly over multiple years. The RECOVER study features negative-unlabeled outcomes, as it includes participants without a history of SARS-CoV-2 infection. These uninfected individuals establish the baseline symptom prevalence and variability in the general population and can thus serve as a valuable control group for studying LC, provided that the statistical model is capable of appropriately accommodating negative-unlabeled data. The combination of partial labeling of LC status, high-dimensional heterogeneous symptom data, and irregular follow-up schedules in RECOVER presents significant analytic challenges that need to be addressed by the statistical methods proposed in this dissertation.

Chapter 1 proposes a sparse Bernoulli mixture model (BMM) with novel parameterization to achieve feature selection. The method can identify symptom-based latent clusters and select a minimally informative set of features. We extend the proposed BMM to leverage negative-unlabeled data to improve the cluster identification. An application to the RECOVER-Adult Cohort data reveals 3 LC indeterminate clusters and 2 LC subphenotypes, as well as 11 crucial symptoms for defining LC.

Chapter 2 extends this framework to longitudinal data by developing a latent Markov model (LMM) with structured transition matrices. The proposed LMM accommodates negative-unlabeled outcomes by incorporating data from uninfected individuals, yielding improved performance compared to approaches that rely solely on infected participants. In addition, the model naturally handles sparse longitudinal data arising from missed visits and staggered enrollment. Chapter 3 applies the proposed LMM to the RECOVER-Adult Cohort data and further adapts the model to include clinical covariates. The method uncovers 5 latent clusters and identified key factors associated with LC persistence or recovery. Collectively, these methods provide a flexible and robust analytic framework for negative-unlabeled data, with applications both within and beyond the context of LC.

Description

Other Available Sources

Research Data

Keywords

feature selection, Long COVID, Markov model, mixture model, SARS-CoV-2, Biostatistics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories