Publication: Statistical Methods for Institution-Scale Science
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Biomedical science of the 21st century is characterized in part by the establishment of large-scale, multi-institutional data collection initiatives, such as the UK Biobank, eMERGE, and All of Us programs. These initiatives provide researchers with a wealth of high-dimensional medical data with which they can ask and answer new scientific questions, but also introduce unique data analysis challenges, specifically concerning data privacy, integration of data from diverse sources, and the careful modeling of complex, longitudinal healthcare data. This thesis aims to address each of these challenges in turn through the development of statistical methods and theory. The first chapter, Multi-task learning with summary statistics, describes a general framework for fitting high-dimensional linear models using publicly summary statistics from GWAS. The second, titled Fast and robust invariant generalized linear models, outlines a nonconvex optimization algorithm for efficiently computing invariant regression coefficients from multi-environment data. The final chapter, Latent factor point processes for patient representation with electronic health records, describes a new nonparametric point process model for longitudinal health records data, and provides a representation learning algorithm that captures meaningful clinical signal under this model.