Publication: Learning and Benchmarking Temporal Representations in Claims Data: Telehealth Trajectories in Buprenorphine Treatment for Opioid Use Disorder
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
The growing availability of longitudinal administrative claims data has created unprecedented opportunities to study healthcare utilization, treatment engagement, and patient trajectories in real-world settings. Yet despite this abundance of real-world data, extracting meaningful temporal insights remains challenging because claims records are sparse, irregularly timed, and highly sensitive to operational definitions. These challenges are particularly relevant when studying clinical patterns during buprenorphine treatment for opioid use disorder (OUD), where healthcare interactions are recorded as sparse longitudinal events, yet it remains unclear how different temporal representation strategies characterize these patterns and whether they identify meaningful utilization phenotypes. To address this gap, this work uses telehealth utilization during buprenorphine treatment for OUD as a case study to evaluate how different temporal representation strategies characterize sparse longitudinal claims data. The primary objective was to compare whether rule-based definitions, deterministic timing features, linear embeddings, latent trajectory models, and recurrent neural network embeddings recover distinct telehealth-use phenotypes beyond conventional early-versus-late exposure definitions. A secondary objective was to assess whether these phenotypes were robust to alternative buprenorphine episode definitions. Buprenorphine treatment episodes were constructed from prescription claims using allowable refill gaps of 7, 14, 30, and 60 days. Telehealth encounters were linked to treatment episodes and
analyzed primarily within the COVID-19-era study window. Five approaches were compared: binary and early-versus-late telehealth definitions, normalized timing embeddings, sparse matrix embeddings using truncated singular value decomposition, latent class growth analysis (LCGA), and gated recurrent unit (GRU) autoencoder embeddings. Methods were benchmarked using internal clustering metrics, bootstrap stability, trajectory diagnostics, early-threshold sensitivity analyses, and cross-gap reassignment patterns. Across methods, telehealth trajectories were sparse and heterogeneous. Simple early-versus-late definitions mainly captured time to first telehealth use and could not distinguish transient early use from sustained longitudinal engagement. SVD and normalized embeddings provided limited additional resolution. LCGA identified four interpretable timing–intensity phenotypes and showed the strongest cross-gap preservation, while GRU embeddings achieved stronger quantitative clustering stability but largely collapsed trajectories into broader low- versus high-utilization patterns. Overall, this case study shows that when studying sparse, irregular, and definition-sensitive claims data, temporal representation choice materially shapes the resulting phenotype space. More complex models do not automatically produce more clinically interpretable subgroups; instead, benchmarking should explicitly evaluate stability, interpretability, robustness to episode construction, and the specific behavioral dimensions each method captures.