Publication:

Causal Inference Using External Comparators: Identification, Estimation, and Applications

Loading...
Thumbnail Image

Date

2026-06-05

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Ung, Lawson. 2026. Causal Inference Using External Comparators: Identification, Estimation, and Applications. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

There is increasing clinical, academic, and regulatory interest in combining information from experimental studies, including randomized and single-group trials, with external information drawn from experimental or observational data sources. Such efforts are usually motivated by the desire to compare treatments evaluated in different studies, sometimes described as external comparator analyses. The overarching goal of this dissertation is to formalize external comparator analyses using the language of potential, or counterfactual, outcomes, with a view to develop rigorous causal and statistical methods for their conduct. The methods are accompanied by practical demonstrations to answer clinically important causal questions under assumptions that are plausible in light of subject matter expertise within their given domains.

In Chapter 1, we begin by describing the causal problem in the non-failure time setting with a time-fixed treatment and a binary, count, or continuous outcome. We introduce a family of basic study templates and elaborate identification strategies for potential outcome means and average treatment effects, focusing on causal contrasts that would not be identifiable by using information on any one of the data sources in isolation. We argue that external comparator analyses inherit concerns relevant to the study of causation in single-source studies, as well as the related literature on combining information (e.g., generalizability and transportability methods). However, external comparator analyses merit consideration as a separate class of causal problems because they differ in terms of their scientific motivations, target population definitions, sampling considerations, data structures, and identifiability conditions. The mathematical formalism and notation introduced in this chapter is adopted throughout the dissertation.

In Chapter 2, we focus on a subclass of external comparator analyses delineated by two distinct, but related, identification strategies, namely identification under transportability in mean and in effect measure. We propose semiparametric efficient estimators, derive their asymptotic properties -- including their robustness to both model misspecification and slower rates of convergence for some nuisance function models -- and use simulation to evaluate their finite sample performance compared to estimators based only on outcome modeling and weighting. We demonstrate the methods by combining the randomized ACCEPT (NCT00454584) and PHOENIX 1 (NCT00267969) trials to evaluate biologic agents in autoimmune disease, specifically moderate-to-severe plaque psoriasis.

In Chapter 3, we extend identification and estimation approaches to the discrete causal survival analysis setting, addressing analytical challenges presented by time-to-event outcomes, censoring, and competing events. We address identification under transportability in mean and in effect measure, and derive discrete survival estimators based on parametric standardization (g-formula) and weighting. We apply the methods to the combined analysis of two landmark cardiology trials, RE-LY (NCT00262600) and ENGAGE AF-TIMI 48 (NCT00781391), to compare the direct-acting oral anticoagulants dabigatran and edoxaban. We specifically address their efficacy in terms of stroke and systemic embolism, and their safety in terms of major bleeding events.

Finally, in Chapter 4, we apply the methods developed in the first three chapters by conducting a nested external comparator analysis, transporting its findings to a target population of interest. Specifically, we combine two legacy lung cancer screening trials, the National Lung Screening Trial (NLST, NCT00047385) and the Prostate, Lung, Colorectal, and Ovarian Cancer Screening Trial (PLCO, NCT00339495), to evaluate lung cancer screening strategies in a US representative target population eligible for screening. We focus specifically on the causal contrast comparing screening based on low-dose computed tomography versus no formal screening recommendation, which has never been studied in the context of a US-based trial, and would no longer be ethically permissible due to loss of clinical equipoise.

Collectively, the methods developed and applied in this dissertation may serve to reposition external comparator analyses as a distinct class of causal problems, rather than a mere series of ad hoc statistical adjustments with no clear causal motivation. Framing external comparator analyses in the causal formalism of potential outcomes may allow for the assumptions of these emerging causal problems to be openly communicated and reasoned about in practical applications.

Description

Other Available Sources

Research Data

Keywords

causal inference, combining information, external comparators, identification strategies, robust estimation, transportability, Epidemiology, Public health, Statistics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories