Publication:

Development and Validation of an Interpretable Risk Prediction Model for Stillbirth using Machine Learning

Loading...
Thumbnail Image

Date

2026-05-15

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Dantuluri, Neha Varma. 2026. Development and Validation of an Interpretable Risk Prediction Model for Stillbirth using Machine Learning. Masters Thesis, Harvard Medical School.

Abstract

Background: Stillbirth, defined as fetal death at ≥20 weeks gestation, affects approximately 1 in 175 deliveries in the United States. Disparities linked to social determinants of health (SDOH) contribute an estimated 60% excess risk in historically marginalized populations. Despite growing interest in applying machine learning (ML) to obstetric risk stratification, existing models are limited by retrospective single-site designs, absence of time-series features derived from longitudinal clinical measurements, and lack of external validation. This study developed and externally validated an explainable ML model for stillbirth prediction using features available by 20 weeks of gestation.

Methods: We analyzed 162,374 pregnancies (635 stillbirths; 0.4%) from the Mass General Brigham electronic health record (EHR) system from 2015–2025. Four models were evaluated: Logistic Regression, ElasticNet, XGBoost with random under-sampling (RUS), and Random Forest with RUS. Features included summary statistics derived from repeated prenatal vital sign measurements (blood pressure, BMI, weight trajectories), clinical comorbidities, parity, and SDOH variables including the Area Deprivation Index (ADI) and insurance type. Missing data were addressed using Multiple Imputation by Chained Equations (MICE). External validation was conducted on the nuMoM2b cohort (8,588 nulliparous pregnancies; 45 stillbirths). Model discrimination was compared via DeLong's test with Bonferroni correction; clinical utility was assessed by decision curve analysis (DCA).

Results: Random Forest + RUS demonstrated the best overall performance and was selected as the final model. Internal validation yielded AUROC 0.797 (95% CI: 0.764–0.832), sensitivity 68.9% (63.8–74.4%), and specificity 71.8% (70.2–73.8%). External validation on nuMoM2b confirmed aggregate-level generalizability with AUROC 0.789 (95% CI: 0.683–0.879), though subgroup performance was substantially attenuated due to structural data differences between cohorts. DCA demonstrated net clinical benefit across probability thresholds of 0.1–2%. SHAP analysis identified weight variability during pregnancy, prenatal visit frequency, and ADI rank as the top predictors.

Conclusions: We present an externally validated, EHR-based ML model for early stillbirth risk prediction that integrates SDOH, clinical, and time-series features to achieve clinically meaningful discrimination. The model demonstrated stable internal performance across racial, ethnic, and insurance-defined subgroups, though external subgroup performance was limited by cohort-specific data constraints. Post-hoc recalibration and prospective validation in diverse populations are necessary before clinical deployment.

Description

Other Available Sources

Research Data

Keywords

Machine Learning, Stillbirth, Bioinformatics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories