Publication:

Optimal Inference in High-Dimensional Structured Models

Loading...
Thumbnail Image

Date

2026-05-11

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Sankaranarayanan, Madhav. 2026. Optimal Inference in High-Dimensional Structured Models. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Across many disciplines, there is a need for the analysis and interpretation of large-scale datasets. While there are many methodologies for high-dimensional regimes that are utilized in practice, the theoretical guarantees of these methodologies can be understudied, even unsubstantiated. In this dissertation, we explore different inferential settings and investigate theoretical gaps pertaining to optimality. We tackle these theoretical gaps using stylized statistical setups. Each chapter focuses on a different inferential motivation, but all fall under the larger pantheon of high-dimensional inference. Chapters 1 and 2 provide novel, provably-optimal estimation procedures in two different statistical settings. Chapters 1 and 3 focus on linear regression, providing new results in optimal estimation and optimal detection under unmeasured confounding, respectively. Chapters 2 and 3 are primarily concerned with minimax testing boundaries and the construction of optimal tests. This work lays the groundwork for understanding many facets of high-dimensional inference through a more rigorous lens.

In Chapter 1, we consider statistical inference in high-dimensional regression problems under affine constraints on the parameter space. The theoretical study of this is motivated by the study of genetic determinants of diseases, such as diabetes, using external information from mediating protein expression levels. Specifically, we develop rigorous methods for estimating genetic effects on diabetes-related continuous outcomes when these associations are constrained based on external information about genetic determinants of proteins, and genetic relationships between proteins and the outcome of interest. In this regard, we discuss multiple candidate estimators and study their theoretical properties, sharp large sample optimality, and numerical qualities under a high-dimensional proportional asymptotic framework. Finally, we apply the developed methods to study the genetic determinants of BMI, fasting insulin and HbA1c, leveraging their genetic correlation with protein expression obtained from an external study.

In Chapter 2, we investigate the problem of minimax optimal detection of the number of spikes in lower rank signal plus large Wigner models. In this regard, we categorize sharp conditions for asymptotic power in testing for the number of spikes being $k_0$ versus $k_1$ in terms of $k_1,k_0$ and the spectral gap between the first $k_0$ and last $k_1-k_0$ spiked eigenvalues. Additionally, when spikes are bounded, we elucidate a novel optimal test in the subcritical regime under a Gaussian Orthogonal Ensemble. For the general hypothesis testing problem, we compare the performance of some ubiquitous tests in spike detection problems. We also elucidate an algorithm for estimating the number of spikes in outlying spiked setups.

In Chapter 3, we address the problem of optimal inference under unmeasured confounding. Instrumental Variables (IV) are a ubiquitous class of methods for adjusting potentially endogenous variables in regression models. In this paper, we address the problem of detecting sparse signal in the case of high-dimensional exposures and instruments. We derive sharp detection thresholds for testing, where we detect sparse regression vectors that are bounded away from $0$ in $\ell_2$ norm. We show that this minimax separation rate is a function of the instrumental variable structure, which acts as a sample size inflation'' factor. Our results extend beyond the proportional asymptotic regime, to accurately capture the problem of detection under a many weak IVs'' regime, which is commonly employed in econometrics and genomic literature.

Description

Other Available Sources

Research Data

Keywords

Asymptotic inference, High-dimensional theory, Optimality, Random matrix theory, Regression models, Statistics, Biostatistics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories