Publication: Speech as a Digital Biomarker: A Computational Framework for Severity and Progression in Parkinson's Disease and Cerebellar Ataxia
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Parkinson's disease and cerebellar ataxia are progressive neurodegenerative conditions monitored through clinician-administered rating scales that cannot support the continuous assessment needed to detect subtle change or serve as sensitive clinical trial endpoints. Speech, non-invasively collectible at scale and disrupted early and prominently in both conditions, offers a path toward objective monitoring. This thesis develops and evaluates a computational framework for speech-based severity assessment using longitudinal data from the Neurobooth platform at Massachusetts General Hospital, where standardized diadochokinetic speech tasks and same-day clinical assessments (MDS-UPDRS Part III; BARS) are collected within a single clinic visit for 163 PD patients, 217 ataxia patients, and 87 healthy controls across 921 sessions.
Three modeling approaches are evaluated across seven cross-sectional prediction targets: a handcrafted acoustic feature pipeline with classical machine learning (LASSO, XGBoost, SVM); a deep learning framework based on mel spectrogram gradient representations and ResNet-18; and frozen transformer-based transfer learning using wav2vec2 and WavLM. From diadochokinetic speech tasks alone, both disease cohorts are separable from healthy controls (AUC 0.824 for PD, 0.857 for ataxia) and from each other (AUC 0.863). For ataxia severity, BARS total R² reaches 0.621 (ResNet-18 with strong L2 regularization) and 0.620 (WavLM Transfer), competitive with the published benchmark on a separate cohort, and WavLM achieves the best BARS speech result (R² = 0.601). A cross-disease biomarker analysis identifies 62 shared acoustic features, 66 PD-specific, and 85 ataxia-specific, characterizing the acoustic divergence between hypokinetic and ataxic dysarthria.
Post-hoc sensitivity analyses show that deep learning models trained on cross-sectional severity produce statistically significant upward drift in predictions across patient sessions (Cohen's d = 0.249-0.418, all four targets) despite no longitudinal supervision. WavLM Transfer additionally separates pharmacological state at the group level in Mild PD without medication supervision: OFF-state sessions receive significantly higher predicted severity than ON-state sessions (p = 0.003, rank-biserial r = -0.235). Taken together, these findings support the further development of speech-based measurement as a quantitative tool for remote assessment, longitudinal monitoring, and clinical trial outcome measurement in neurodegenerative disease.