Publication: Policy Learning under Interference and Adaptively Collected Data
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Artificial intelligence has progressed significantly in the recent decades thanks to the wealth of available data and computational power. One area that has made great strides is policy learning, or the development of data-driven algorithms that learn optimal decision-making rules, for some objective of interest. These algorithms are applied across industry, driving the technology behind soccer-playing robots to large language models that give more realistic natural language responses. Despite its widespread use, there remain methodological gaps for effective use in public health. This dissertation addresses this gap proposing novel methods for policy learning under interference and adaptively collected data settings.
In Chapter 1, we propose an approach for optimal policy learning algorithm under bipartite network interference (BNI). BNI differs from classical causal inference settings in two ways: (i) the unit treated differs from the unit on which outcomes are observed, and (ii) the treatment exposure may spillover across multiple outcome units. We propose the first policy learning methods under BNI, providing statistical inference procedures and optimality guarantees. We then employ this methodology to study the optimal cost-constrained environmental policy for flue-gas desulfurization units, or scrubbers, onto coal fired power plants.
In Chapter 2, we build on the results from Chapter 1 to consider a different lens of policymaking under BNI. Numerous studies have shown that marginalized subgroups, such as low income or older communities, may experience disproportionate health burden from environmental exposures. We propose a novel fair policy learning methodology under BNI that minimizes welfare disparity between subgroups under the constraint of Pareto optimality. By framing fairness objectives under Pareto optimality, we ensure that subgroups cannot gain without inducing harm upon another group, staying in line with the principle of `first do no harm'. We employ this method to study fair, cost-effective scrubber allocations in the U.S. that balances mortality risk among high and low poverty subgroups, analyzing the tradeoffs needed to achieve fairness in environmental policymaking.
In Chapter 3, we shift gears to consider adaptively collected data settings. Despite its extensive use across mobile health and online advertising, many statistical questions on nonparametric estimation and inference remain. To that end, we study the problem of nonparametric regression under adaptively collected data settings, characterizing minimax optimal $L^2$ and $L^\infty$ rates. We then propose estimators which achieve minimax optimality, while permitting asymptotic normality and uniform inference under `high exploitation', regret minimizing settings. As an application, we extend the pseudo-outcome framework to permit conditional average treatment effect estimation and inference.