Publication:

Policy Learning under Interference and Adaptively Collected Data

Loading...
Thumbnail Image

Date

2026-02-27

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Kim, Raphael. 2026. Policy Learning under Interference and Adaptively Collected Data. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Artificial intelligence has progressed significantly in the recent decades thanks to the wealth of available data and computational power. One area that has made great strides is policy learning, or the development of data-driven algorithms that learn optimal decision-making rules, for some objective of interest. These algorithms are applied across industry, driving the technology behind soccer-playing robots to large language models that give more realistic natural language responses. Despite its widespread use, there remain methodological gaps for effective use in public health. This dissertation addresses this gap proposing novel methods for policy learning under interference and adaptively collected data settings.

In Chapter 1, we propose an approach for optimal policy learning algorithm under bipartite network interference (BNI). BNI differs from classical causal inference settings in two ways: (i) the unit treated differs from the unit on which outcomes are observed, and (ii) the treatment exposure may spillover across multiple outcome units. We propose the first policy learning methods under BNI, providing statistical inference procedures and optimality guarantees. We then employ this methodology to study the optimal cost-constrained environmental policy for flue-gas desulfurization units, or scrubbers, onto coal fired power plants.

In Chapter 2, we build on the results from Chapter 1 to consider a different lens of policymaking under BNI. Numerous studies have shown that marginalized subgroups, such as low income or older communities, may experience disproportionate health burden from environmental exposures. We propose a novel fair policy learning methodology under BNI that minimizes welfare disparity between subgroups under the constraint of Pareto optimality. By framing fairness objectives under Pareto optimality, we ensure that subgroups cannot gain without inducing harm upon another group, staying in line with the principle of `first do no harm'. We employ this method to study fair, cost-effective scrubber allocations in the U.S. that balances mortality risk among high and low poverty subgroups, analyzing the tradeoffs needed to achieve fairness in environmental policymaking.

In Chapter 3, we shift gears to consider adaptively collected data settings. Despite its extensive use across mobile health and online advertising, many statistical questions on nonparametric estimation and inference remain. To that end, we study the problem of nonparametric regression under adaptively collected data settings, characterizing minimax optimal $L^2$ and $L^\infty$ rates. We then propose estimators which achieve minimax optimality, while permitting asymptotic normality and uniform inference under `high exploitation', regret minimizing settings. As an application, we extend the pseudo-outcome framework to permit conditional average treatment effect estimation and inference.

Description

Other Available Sources

Research Data

Keywords

Biostatistics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories