Publication:

Models of Human Decision-Making for Planning Digital Interventions Using Reinforcement Learning

Loading...
Thumbnail Image

Date

2026-01-06

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Nofshin, Eura. 2026. Models of Human Decision-Making for Planning Digital Interventions Using Reinforcement Learning. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Digital interventions support people in sustaining effort toward their long-term goals, yet the user-level data available to personalize these interventions is scarce and noisy. This dissertation studies how reinforcement learning (RL) systems can make high-quality decisions under these data limitations by modeling the human user.

The first part of the thesis focuses on what human model to use. Drawing from behavioral science, I formalize Behavior Model RL (BMRL), a two-agent framework in which the human is represented as a sequential decision-maker with potentially maladapted Markov Decision Process (MDP) parameters, and the AI intervenes on these parameters to help reach their goal. BMRL requires us to specify how the AI agent models the other human agent. In practice, we must simplify our human models to support online learning for the AI, even knowing that these assumptions do not fully reflect real users. For instance, I introduce a simple, computationally tractable model of a user's goal-directed behavior in digital settings. I also point out that, as a field, RL simplifies its agents' discount models; we use exponential discounting to model agents (including human ones), despite evidence from psychology that humans discount hyperbolically. I examine how such modeling choices affect an AI’s ability to learn policies online, and I develop tools to understand whether these simplified models can generalize to more complex human behaviors.

The second part of the thesis examines how to learn human models online. I study the bias-variance trade-offs that arise when personalizing to individuals with limited data, and propose algorithms that manage model complexity over time. One method learns how model complexity should evolve by transferring "kernel evolution"' trajectories from prior users to new users, enabling rapid and stable online model selection in Gaussian Process regression. Another method increases the size of the AI's state space model of the human as more data is collected.

Together, these contributions provide a computational foundation for embedding behavioral insights into human models so that we can plan digital interventions using RL. They illustrate that the right inductive biases can enable AI systems to plan effective, interpretable, and personalized support for behavior change.

Description

Other Available Sources

Research Data

Keywords

Computer science

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories