Publication:

Exact Asymptotics in the Evaluation of Sequential Experiments and High-Dimensional Models

Loading...
Thumbnail Image

Open/View Files

Date

2026-05-18

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Wang, Longlin. 2026. Exact Asymptotics in the Evaluation of Sequential Experiments and High-Dimensional Models. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Modern statistical learning procedures and sequential decision-making systems are increasingly used in settings where reliable performance evaluation and risk assessment are essential. Yet many classical guarantees are organized around conservative minimax bounds, fixed-target assumptions, or asymptotic regimes that can obscure the behavior of the adaptive and high-dimensional algorithms used in practice. This dissertation develops exact asymptotic theory for three related problems: simultaneous learning and evaluation in reinforcement learning, risk characterization in high-dimensional transfer learning, and variable selection under correlated designs.

The first part of the dissertation studies off-policy evaluation in reinforcement learning. Standard off-policy evaluation treats the target policy as exogenously fixed, whereas practical workflows often train and evaluate a policy using the same trajectory. This data reuse induces adaptive evaluation bias and can invalidate standard uncertainty quantification. To address this problem, we extend the CRAM methodology to Markov Decision Processes, enabling simultaneous learning and evaluation of a target policy along a single adaptively sampled trajectory without data-inefficient sample splitting. Under an algorithmic stability condition that controls the discrepancy between consecutive learned policies, we establish consistency and asymptotic normality of the CRAM estimator and construct a fully data-driven variance estimator. We further verify that this stability condition is compatible with several broad classes of reinforcement learning algorithms, including policy gradient methods, deterministic offline algorithms, and entropy-regularized procedures.

The second part develops exact risk characterizations for high-dimensional transfer learning under distribution shifts. While multi-stage estimators are commonly used to combine auxiliary source data with a target environment, their theoretical performance has often been described through upper bounds that suppress constants and procedure-specific differences. We introduce Multi-Environment Generalized Long Approximate Message Passing (GLAMP), an AMP framework that accommodates matrix-valued iterates, non-separable denoising functions, multiple design matrices, and anisotropic covariance structures. Through a rigorous state evolution analysis in the proportional asymptotic regime, we derive exact mean-squared error characterizations for three representative Lasso-based transfer learning methods: the Stacked Lasso, the Model Averaging Estimator, and the Second-Step Estimator.

The final part provides a detailed theoretical comparison of variable selection methods in sparse linear models with blockwise correlated designs. Moving beyond exact support recovery, we evaluate performance through the expected Hamming error under a rare-and-weak signal model. We derive exact convergence exponents and corresponding phase diagrams for six widely used methods: Lasso, Elastic net, SCAD, thresholded Lasso, forward selection, and forward backward selection. These phase transitions delineate regions of Exact Recovery, Almost Full Recovery, and No Recovery, and they quantify how correlation structure affects the relative performance of different selection procedures. In particular, the analysis shows how non-convex penalties such as SCAD mitigate negative-correlation signal cancellation, and how post-processing and backward elimination can substantially reduce selection error in the model studied here.

Collectively, this dissertation advances the theoretical evaluation of complex, data-driven statistical procedures. By moving beyond conservative bounds toward exact operational limits, it develops mathematically tractable tools for algorithm comparison, hyperparameter tuning, and reliable deployment in adaptive and high-dimensional settings.

Description

Other Available Sources

Research Data

Keywords

Statistics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories