Publication: Optimal Inference in High-Dimensional Structured Models
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Across many disciplines, there is a need for the analysis and interpretation of large-scale datasets. While there are many methodologies for high-dimensional regimes that are utilized in practice, the theoretical guarantees of these methodologies can be understudied, even unsubstantiated. In this dissertation, we explore different inferential settings and investigate theoretical gaps pertaining to optimality. We tackle these theoretical gaps using stylized statistical setups. Each chapter focuses on a different inferential motivation, but all fall under the larger pantheon of high-dimensional inference. Chapters 1 and 2 provide novel, provably-optimal estimation procedures in two different statistical settings. Chapters 1 and 3 focus on linear regression, providing new results in optimal estimation and optimal detection under unmeasured confounding, respectively. Chapters 2 and 3 are primarily concerned with minimax testing boundaries and the construction of optimal tests. This work lays the groundwork for understanding many facets of high-dimensional inference through a more rigorous lens.
In Chapter 1, we consider statistical inference in high-dimensional regression problems under affine constraints on the parameter space. The theoretical study of this is motivated by the study of genetic determinants of diseases, such as diabetes, using external information from mediating protein expression levels. Specifically, we develop rigorous methods for estimating genetic effects on diabetes-related continuous outcomes when these associations are constrained based on external information about genetic determinants of proteins, and genetic relationships between proteins and the outcome of interest. In this regard, we discuss multiple candidate estimators and study their theoretical properties, sharp large sample optimality, and numerical qualities under a high-dimensional proportional asymptotic framework. Finally, we apply the developed methods to study the genetic determinants of BMI, fasting insulin and HbA1c, leveraging their genetic correlation with protein expression obtained from an external study.
In Chapter 2, we investigate the problem of minimax optimal detection of the number of spikes in lower rank signal plus large Wigner models. In this regard, we categorize sharp conditions for asymptotic power in testing for the number of spikes being $k_0$ versus $k_1$ in terms of $k_1,k_0$ and the spectral gap between the first $k_0$ and last $k_1-k_0$ spiked eigenvalues. Additionally, when spikes are bounded, we elucidate a novel optimal test in the subcritical regime under a Gaussian Orthogonal Ensemble. For the general hypothesis testing problem, we compare the performance of some ubiquitous tests in spike detection problems. We also elucidate an algorithm for estimating the number of spikes in outlying spiked setups.
In Chapter 3, we address the problem of optimal inference under unmeasured confounding. Instrumental Variables (IV) are a ubiquitous class of methods for adjusting potentially endogenous variables in regression models. In this paper, we address the problem of detecting sparse signal in the case of high-dimensional exposures and instruments. We derive sharp detection thresholds for testing, where we detect sparse regression vectors that are bounded away from $0$ in $\ell_2$ norm. We show that this minimax separation rate is a function of the instrumental variable structure, which acts as a sample size inflation'' factor. Our results extend beyond the proportional asymptotic regime, to accurately capture the problem of detection under a many weak IVs'' regime, which is commonly employed in econometrics and genomic literature.