Publication: Genetic and Proteomic Factors Underlying Complex Human Diseases
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
The rapid growth of genomic data over the past two decades has brought the biomedical field closer to precision medicine, which aims to develop data-driven and tailored approaches for disease prevention and treatment. Population-scale biobanks have played an instrumental role in generating vast amounts of human genetic sequencing data, empowering extensive investigations into the genetics underlying complex diseases. Notably, biobanks have facilitated thousands of genome-wide association studies (GWAS). These studies use genotype information from large cohorts to identify common genetic variants associated with phenotypes of interest, and have provided foundational insights into the genetic basis of numerous complex diseases and medically-relevant traits. However, fulfilling the vision of precision medicine requires deeper understanding of genetics in diverse populations and within broader biological contexts.
First, most GWAS have been conducted in populations of European ancestry, which has not only limited GWAS discoveries but also the clinical translation of GWAS findings. Associations identified in GWAS are frequently utilized to construct polygenic risk scores (PRS), estimates of individuals’ genetic risk that have been incorporated into clinical risk models for some complex diseases. Imbalances in the representation of different populations in biobanks and GWAS have precipitated the development of many PRS with lower predictive accuracy in populations of non-European ancestries.
Second, while genes provide the blueprint for biology, they are several steps removed from the molecular mechanisms driving disease. Biomarkers from other -omic data, such as proteomics, may help bridge the gap between the genome and disease etiology, providing a more comprehensive view of underlying causal pathways as well as the effects of non-genetic factors on disease processes. Emerging datasets of plasma proteomics from population biobanks have highlighted the great potential of this data for precision medicine, but challenges remain in delineating the interplay between the genome, proteome, environment, and disease.
In Chapter 1, I conduct a multi-ancestry GWAS meta-analysis using data from 18 biobanks to characterize the genetic architecture of asthma, a heterogeneous and multifactorial disease with variable prevalence rates. I identify 49 novel associations, and demonstrate the value of integrating data from many diverse cohorts to assess shared genetic risk across ancestries, biobanks, and disease subtypes, as well as improve polygenic risk prediction.
In Chapter 2, I develop PRS trained on multi-ancestry and multi-biobank data with up to 750,000 participants for 32 complex diseases and traits. By evaluating the prediction performance of these models across diverse ancestry groups, I elucidate strategies for constructing the optimal PRS depending on ancestry, method, and genetic architecture.
In Chapter 3, I dissect the mechanisms driving thousands of associations between 2,935 plasma proteins and onset of 23 age-related diseases. By integrating disease and protein GWAS data in a causal inference approach, I identify a subset of proteomic biomarkers that play causal roles in disease development. I also demonstrate that a large proportion of the circulating proteome is associated with smoking, and develop a proteomic score that captures smoking behavior and history.
Together, these studies leverage multi-omic data to expand our understanding of the factors underlying disease risk, with the ultimate goal of advancing precision medicine.