Kovacheva, Vesela PManrai, Arjun KDantuluri, Neha Varma2026-05-1920262026-05-152026Dantuluri, Neha Varma. 2026. Development and Validation of an Interpretable Risk Prediction Model for Stillbirth using Machine Learning. Masters Thesis, Harvard Medical School.32696095https://dash.harvard.edu/handle/1/42738339Background: Stillbirth, defined as fetal death at ≥20 weeks gestation, affects approximately 1 in 175 deliveries in the United States. Disparities linked to social determinants of health (SDOH) contribute an estimated 60% excess risk in historically marginalized populations. Despite growing interest in applying machine learning (ML) to obstetric risk stratification, existing models are limited by retrospective single-site designs, absence of time-series features derived from longitudinal clinical measurements, and lack of external validation. This study developed and externally validated an explainable ML model for stillbirth prediction using features available by 20 weeks of gestation. Methods: We analyzed 162,374 pregnancies (635 stillbirths; 0.4%) from the Mass General Brigham electronic health record (EHR) system from 2015–2025. Four models were evaluated: Logistic Regression, ElasticNet, XGBoost with random under-sampling (RUS), and Random Forest with RUS. Features included summary statistics derived from repeated prenatal vital sign measurements (blood pressure, BMI, weight trajectories), clinical comorbidities, parity, and SDOH variables including the Area Deprivation Index (ADI) and insurance type. Missing data were addressed using Multiple Imputation by Chained Equations (MICE). External validation was conducted on the nuMoM2b cohort (8,588 nulliparous pregnancies; 45 stillbirths). Model discrimination was compared via DeLong's test with Bonferroni correction; clinical utility was assessed by decision curve analysis (DCA). Results: Random Forest + RUS demonstrated the best overall performance and was selected as the final model. Internal validation yielded AUROC 0.797 (95% CI: 0.764–0.832), sensitivity 68.9% (63.8–74.4%), and specificity 71.8% (70.2–73.8%). External validation on nuMoM2b confirmed aggregate-level generalizability with AUROC 0.789 (95% CI: 0.683–0.879), though subgroup performance was substantially attenuated due to structural data differences between cohorts. DCA demonstrated net clinical benefit across probability thresholds of 0.1–2%. SHAP analysis identified weight variability during pregnancy, prenatal visit frequency, and ADI rank as the top predictors. Conclusions: We present an externally validated, EHR-based ML model for early stillbirth risk prediction that integrates SDOH, clinical, and time-series features to achieve clinically meaningful discrimination. The model demonstrated stable internal performance across racial, ethnic, and insurance-defined subgroups, though external subgroup performance was limited by cohort-specific data constraints. Post-hoc recalibration and prospective validation in diverse populations are necessary before clinical deployment.application/pdfenMachine LearningStillbirthBioinformaticsDevelopment and Validation of an Interpretable Risk Prediction Model for Stillbirth using Machine LearningThesis or Dissertation2026-05-190009-0003-7986-1523