Publication: Real-World Evidence in Oncology Using Reference Trial Emulation Across Multiple Electronic Health Record Databases
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Background: Oncology specialty electronic health record (EHR) databases are increasingly used to generate real-world evidence (RWE) in prognostic modeling and comparative effectiveness research. These data sources offer the potential to complement randomized clinical trials (RCTs) by reflecting routine clinical practice and including more heterogeneous patient populations. However, their use remains challenged by incomplete capture of key prognostic variables, missing data, and heterogeneity across databases. As a result, it remains uncertain under what conditions oncology EHR data are sufficiently complete and reliable to support valid prognostic assessment and causal inference. Benchmarking real-world analyses against RCTs has been proposed as an approach for assessing fitness-for-purpose of a given real-world data source for specific, closely related comparative effectiveness questions. Objectives: The objectives of this body of work were to: (1) evaluate the generalizability and performance of a computable prognostic score for overall survival (OS) across multiple oncology specialty EHR-derived databases; and (2)/(3) explore the extent to which multiple oncology specialty EHR-derived data can replicate OS treatment effects observed in two oncology RCTs. Methods: First, we evaluated the performance of the Real-wOrld PROgnostic (ROPRO) score, a multivariable prognostic model derived from EHR-derived data, across four independent oncology EHR databases covering multiple cancer types and disease settings. ROPRO scores were computed using routinely collected demographic, clinical, and laboratory variables, with missing covariates imputed using random forest-based methods. OS was modeled using Cox proportional hazards models, and performance was assessed using discrimination and calibration metrics. Second, we conducted two exploratory comparative effectiveness studies emulating the MONARCH-3 and MONALEESA-2 RCTs, which evaluated cyclin-dependent kinase (CDK) 4/6 inhibitors combined with endocrine therapy as first-line treatment for hormone receptor-positive, human epidermal growth factor receptor 2 (HER2)-negative metastatic breast cancer. Across three oncology specialty EHR-derived databases, we aligned eligibility criteria, treatment definitions, and follow-up with the reference trials. OS was the outcome of interest. Missing data were addressed using multiple imputation by chained equations, and confounding was controlled using 1:1 nearest-neighbor propensity-score (PS) matching. Databases achieving covariate balance after PS matching were included in the primary analysis, which included a fixed-effects inverse-variance meta-analytic approach. All analyses were conducted under preregistered protocols and were considered exploratory due to limited statistical power. Results: The ROPRO demonstrated consistent and generalizable prognostic performance across databases and cancer types when all or nearly all variables were available, with moderate-to-good discrimination and overall adequate calibration. Performance deteriorated in a database with limited variable availability, highlighting the sensitivity of prognostic modeling to data completeness. In the MONARCH-3 emulation, two databases met inclusion criteria for the primary analysis, and the pooled OS hazard ratio (HR) was closely aligned with the randomized trial estimate, although database-specific estimates varied substantially in magnitude and direction. In the MONALEESA-2 emulation, only one database achieved covariate balance; the resulting effect estimate was aligned with the trial result, while additional databases were informative only in sensitivity analyses. Across both emulations, pooled estimates generally approximated trial findings, but uncertainty was substantial and individual databases produced heterogeneous results. Conclusions: Collectively, these studies demonstrate that oncology EHR-derived real-world data can, under certain conditions, support valid prediction and approximate randomized trial estimates for OS. However, fitness-for-purpose is highly context specific and depends critically on data completeness, mortality capture, confounder availability, and line-of-therapy curation. The ROPRO score was generalizable across datasets when key variables are well captured, while confidence in real-world comparative effectiveness analyses may be increased through benchmarking against closely related RCTs. These findings underscore the need for pre-specified study design and analysis, transparent reporting, and database- and question-specific diagnostics to support credible prediction and causal inference in oncology RWE. : First, we evaluated the performance of the Real-wOrld PROgnostic (ROPRO) score, a multivariable prognostic model derived from HER-derived data, across four independent oncology EHR databases covering multiple cancer types and disease settings. ROPRO scores were computed using routinely collected demographic, clinical, and laboratory variables, with missing covariates imputed using random forest-based methods. OS was modeled using Cox proportional hazards models, and performance was assessed using discrimination and calibration metrics. Second, we conducted two exploratory comparative effectiveness studies emulating the MONARCH-3 and MONALEESA-2 RCTs, which evaluated CDK4/6 inhibitors combined with endocrine therapy as first-line treatment for hormone receptor-positive, human epidermal growth factor receptor 2 (HER2)-negative metastatic breast cancer. Across three oncology specialty EHR-derived databases, we aligned eligibility criteria, treatment definitions, and follow-up with the reference trials. OS was the outcome of interest. Missing data were addressed using multiple imputation by chained equations, and confounding was controlled using 1:1 nearest-neighbor propensity-score (PS) matching. Databases achieving covariate balance after PS matching were included in the primary analysis, which included a fixed-effects inverse-variance meta-analytic approach. All analyses were conducted under preregistered protocols and were considered exploratory due to limited statistical power. Results: The ROPRO demonstrated consistent and generalizable prognostic performance across databases and cancer types when all or near-all variables were available, with moderate-to-good discrimination and overall adequate calibration. Performance deteriorated in a database with limited variable availability, highlighting the sensitivity of prognostic modeling to data completeness. In the MONARCH-3 emulation, two databases met inclusion criteria for the primary analysis, and the pooled OS hazard ratio (HR) was closely aligned with the randomized trial estimate, although database-specific estimates varied substantially in magnitude and direction. In the MONALEESA-2 emulation, only one database achieved covariate balance; the resulting effect estimate was aligned with the trial result, while additional databases were informative only in sensitivity analyses. Across both emulations, pooled estimates generally approximated trial findings, but uncertainty was substantial and individual databases produced heterogeneous results. Conclusions: Collectively, these studies demonstrate that oncology EHR-derived real-world data can, under certain conditions, support valid prognostic modeling and approximate randomized trial estimates for OS. However, fitness-for-purpose is highly context specific and depends critically on data completeness, mortality capture, confounder availability, and line-of-therapy curation. The ROPRO are generalizable across datasets when key variables are well captured, while comparative effectiveness analyses may benefit from routine benchmarking against closely related RCTs. These findings underscore the need for pre-specified study design and analysis, transparent reporting, and database- and question-specific diagnostics to support credible causal inference and prognostic assessment in oncology RWE.