Publication: Early Detection of Pediatric Growth Disorders: A Deep Learning Approach to Longitudinal Electronic Health Records
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Pediatric growth abnormalities serve as clinical indicators of a broad spectrum of endocrine, genetic, and gastrointestinal conditions, many of which are treatable when identified early, but carry significant long-term consequences when diagnosis is delayed. Despite advances in screening protocols, diagnostic delays remain prevalent, driven by the longitudinal nature of growth-trajectory deviations and cognitive biases. This study developed a transformer classification model, evaluated it across performance metrics, and analyzed how the model reached its predictions for the early detection of growth-related conditions across 24 diagnoses derived from the International Classification of Pediatric Endocrine Diagnosis (ICPED), using electronic health record data from 9,747 pediatric patients drawn from a pediatric primary care dataset. A central aim was to determine whether strong discriminative ability between patients with growth abnormalities and without could be achieved, and, if so, whether the model’s reasoning patterns aligned with clinically valid reasoning. First, a time-to-event analysis was conducted to establish patient eligibility criteria. Next, transformer and logistic regression models were trained and evaluated across multiple feature configurations. Following optimization, the best-performing model achieved a test AUROC of 0.9823. The interpretability analyses revealed a reasoning mechanism inconsistent with clinical norms. The model operated under a guilty-until-proven-innocent heuristic, assigning elevated predicted risk to all patients by default and reducing that risk after accumulating evidence of normal growth and routine healthcare engagement. Routine encounter codes dominated attribution rankings. The results indicated that performance reflected healthcare utilization patterns more than underlying pathology. Lead time calculation quantified how far in advance of formal diagnosis the model identified at-risk patients, yielding a bootstrapped median estimate of 64.10 months. Given the model’s inverted clinical reasoning, these estimates are better interpreted as upper bounds on true pre-diagnostic capabilities rather than as reliable estimates of clinical utility. These findings demonstrate that transformer architectures can perform classification on longitudinal EHR data, while highlighting that high predictive performance is insufficient evidence of clinical validity. The interpretability analysis developed here offers a replicable methodology for auditing EHR-based diagnostic models before clinical deployment.