Publication:

Deployable Online Reinforcement Learning Algorithms for Use-inspired Research in Digital Health

Loading...
Thumbnail Image

Date

2026-05-12

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Ghosh, Susobhan. 2026. Deployable Online Reinforcement Learning Algorithms for Use-inspired Research in Digital Health. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Online Reinforcement Learning (RL) algorithms have emerged as a promising approach for personalizing just-in-time adaptive interventions (JITAIs) in digital health. Yet, their real-world deployment presents challenges that extend well beyond policy optimization. In digital health settings, online RL algorithms must learn under severe data scarcity and noise, operate autonomously within clinical constraints, remain interpretable and reproducible, and support valid post-deployment scientific analysis. This thesis focuses on methodologies to overcome these challenges to successfully design and deploy online RL in digital health.

To support the continual improvement process of iteratively designing, testing, refining, and re-deploying algorithms, this work first details a reproducible workflow for online AI in digital health, structured across algorithm design, system assurance, and post-deployment analysis. Building on this foundational workflow, the thesis details practical algorithmic solutions to the aforementioned challenges of online RL in digital health. Addressing the algorithm design phase, the thesis presents reBandit, a random effects-based online RL algorithm developed and piloted in the MiWaves digital intervention to reduce cannabis use among emerging adults. To overcome severe data constraints, reBandit leverages informative Bayesian priors and random effects to adaptively pool data across the user population, enabling efficient learning while also personalizing to the individual.

To enhance the stability and autonomy of a deployed algorithm, the work outlines a framework for designing algorithm monitoring systems. For post-deployment analyses, the thesis introduces a resampling-based method to evaluate whether a deployed online RL algorithm successfully learned to personalize treatments. Additionally, it investigates the user experience of participants in the MiWaves pilot study, presenting insights on self-awareness, burden, and intervention personalization to inform the next iteration of the intervention and algorithm. Finally, this thesis concludes by summarizing its key contributions and outlining promising directions for future research.

Description

Other Available Sources

Research Data

Keywords

Digital Health, Mobile Health, Multi-armed bandits, Online RL, Reinforcement Learning, Computer science

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories