Publication:

Learning and Benchmarking Temporal Representations in Claims Data: Telehealth Trajectories in Buprenorphine Treatment for Opioid Use Disorder

Loading...
Thumbnail Image

Date

2026-07-06

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Li, Charis. 2026. Learning and Benchmarking Temporal Representations in Claims Data: Telehealth Trajectories in Buprenorphine Treatment for Opioid Use Disorder. Masters Thesis, Harvard Medical School.

Abstract

The growing availability of longitudinal administrative claims data has created unprecedented opportunities to study healthcare utilization, treatment engagement, and patient trajectories in real-world settings. Yet despite this abundance of real-world data, extracting meaningful temporal insights remains challenging because claims records are sparse, irregularly timed, and highly sensitive to operational definitions. These challenges are particularly relevant when studying clinical patterns during buprenorphine treatment for opioid use disorder (OUD), where healthcare interactions are recorded as sparse longitudinal events, yet it remains unclear how different temporal representation strategies characterize these patterns and whether they identify meaningful utilization phenotypes. To address this gap, this work uses telehealth utilization during buprenorphine treatment for OUD as a case study to evaluate how different temporal representation strategies characterize sparse longitudinal claims data. The primary objective was to compare whether rule-based definitions, deterministic timing features, linear embeddings, latent trajectory models, and recurrent neural network embeddings recover distinct telehealth-use phenotypes beyond conventional early-versus-late exposure definitions. A secondary objective was to assess whether these phenotypes were robust to alternative buprenorphine episode definitions. Buprenorphine treatment episodes were constructed from prescription claims using allowable refill gaps of 7, 14, 30, and 60 days. Telehealth encounters were linked to treatment episodes and

analyzed primarily within the COVID-19-era study window. Five approaches were compared: binary and early-versus-late telehealth definitions, normalized timing embeddings, sparse matrix embeddings using truncated singular value decomposition, latent class growth analysis (LCGA), and gated recurrent unit (GRU) autoencoder embeddings. Methods were benchmarked using internal clustering metrics, bootstrap stability, trajectory diagnostics, early-threshold sensitivity analyses, and cross-gap reassignment patterns. Across methods, telehealth trajectories were sparse and heterogeneous. Simple early-versus-late definitions mainly captured time to first telehealth use and could not distinguish transient early use from sustained longitudinal engagement. SVD and normalized embeddings provided limited additional resolution. LCGA identified four interpretable timing–intensity phenotypes and showed the strongest cross-gap preservation, while GRU embeddings achieved stronger quantitative clustering stability but largely collapsed trajectories into broader low- versus high-utilization patterns. Overall, this case study shows that when studying sparse, irregular, and definition-sensitive claims data, temporal representation choice materially shapes the resulting phenotype space. More complex models do not automatically produce more clinically interpretable subgroups; instead, benchmarking should explicitly evaluate stability, interpretability, robustness to episode construction, and the specific behavioral dimensions each method captures.

Description

Other Available Sources

Research Data

Keywords

buprenorphine, claims data, clinical informatics, opioid use disorder, telehealth, trajectory modeling, Bioinformatics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories