Publication:

Causal Inference in Complex Observational Settings with Applications to Electronic Health Record Data

Loading...
Thumbnail Image

Date

2025-05-13

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Xu, Daniel. 2025. Causal Inference in Complex Observational Settings with Applications to Electronic Health Record Data. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Electronic health record (EHR) data serve as a valuable source of real-world evidence for assessing treatment effects, offering rich, longitudinal patient data that can enhance clinical research, inform healthcare decisions, and ultimately improve patient outcomes. However, leveraging these complex data sources for reliable statistical inference remains challenging. For example, one common difficulty is the limited availability of readily validated clinical outcomes in EHR data, which often requires labor-intensive manual annotation through chart review. A related challenge arises when studies span multiple institutions, where differences in patient populations can introduce heterogeneity and potential biases. These complications further exacerbate the fundamental challenge of using observational data to draw valid causal conclusions. In this dissertation, we propose novel statistical methods for causal inference that address certain complexities relevant to clinical studies leveraging EHR data. In particular, we focus on two settings of interest: (1) semi-supervised learning, where outcome labels are scarce, and (2) transfer learning, where data must be integrated across diverse populations. We hope that the contributions of this dissertation offer meaningful steps toward unlocking the full potential of EHR data for clinical and translational research.

In Chapter 1, we address the problem of estimating treatment effects in a semi-supervised setting where labeled outcomes are sparse but unlabeled data are abundant. We develop a semi-supervised calibration method that leverages a subsample of labeled outcomes to calibrate inferred outcomes, ensuring that the downstream treatment effect estimator remains consistent despite potential errors in outcome imputation. Unlike most traditional semi-supervised methods, we allow the labeling mechanism to depend on the observed data, rather than assuming it is completely random. This problem is analogous to estimating mean outcomes in longitudinal studies with monotone missingness, and we show that our proposed estimator is asymptotically equivalent to the augmented inverse probability weighting (AIPW) estimator when a consistent estimate of the labeling propensity score is available. The estimator is multiply robust and locally semiparametric efficient. We also demonstrate improved finite-sample efficiency in semi-supervised settings, owing to an effective normalization of an implicit augmentation term. The finite-sample performance is evaluated through simulations, and we illustrate the method in a case study comparing the effectiveness of two anti-TNF therapies on remission outcomes in patients with rheumatoid arthritis.

In Chapter 2, we consider the problem of causal mediation analysis in a semi-supervised setting. Causal mediation analysis is a fundamental tool for understanding how treatments or exposures affect outcomes through intermediate variables. However, existing methods lack theoretical and practical guarantees in the presence of substantial missingness in the outcome variable. To address this, we propose a robust and efficient method for the semi-supervised estimation of natural direct and indirect effects, accommodating settings both with and without surrogate outcomes. Our approach extends the double machine learning framework by constructing multiply robust one-step estimators, which enable the use of flexible machine learning algorithms for estimating nuisance functions. We show that the proposed estimators are asymptotically normal and locally minimax optimal under semi-supervised models that do not assume the distribution of the non-missing variables to be known. Through extensive simulations, we demonstrate that the estimators remain unbiased, yield valid inference, and achieve substantial efficiency gains by leveraging predictive surrogates, even under complex data-generating mechanisms. We illustrate the method using EHR data to examine whether racial disparities in Alzheimer’s disease–related outcomes are mediated through cardiovascular comorbidities.

In Chapter 3, we explore the problem of estimating low-dimensional functionals in a target population using data from both source and target domains, where the two distributions are only weakly aligned. This setting arises frequently in practice when performing transfer learning, yet remains underexplored in the context of functional estimation. We focus on three canonical functionals relevant to statistics and causal inference: the quadratic regression functional, the expected conditional covariance, and the mean response in missing data models. Rather than making strong parametric assumptions on the source and target distributions, we model weak alignment by assuming that the differences in conditional mean functions across domains are bounded in L2 norm. For each functional, we propose two estimation strategies: (1) an influence function–based estimator that employs existing minimax-optimal transfer learning methods for nuisance estimation, and (2) a functional-level confidence thresholding estimator that selects between source- and target-based estimators. We derive non-asymptotic upper bounds on the estimation risk in mean absolute error and show that both strategies achieve faster rates than target-only estimators, provided that posterior drift between domains is sufficiently small. When functionals involve two nuisances, we demonstrate that some of the proposed estimators are able to outperform the minimax-optimal target-only estimator as long as the degree of posterior drift in either of the two nuisance functions is sufficiently small, which we refer to as distributional shift double robustness. Furthermore, we demonstrate that under additional assumptions -- for example, by directly bounding the separation between the functional evaluated at source and target distributions -- the functional-level confidence thresholding estimator can attain minimax-optimal rates up to log factors.

Description

Other Available Sources

Research Data

Keywords

Causal Inference, Electronic Health Records, Semi-Supervised Learning, Transfer Learning, Biostatistics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories