Publication: Methods for Multi-ancestry Fine-mapping and Estimation of Frequency-dependent Genetic Architectures
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
In genome-wide association studies (GWAS), each mutation in the genome is tested for association with a disease or trait of interest, one at a time, across hundreds to millions of individuals. GWAS have demonstrated great utility for elucidating human biology and identifying disease-related genes which can be targeted by new drugs. However, the overwhelming majority of GWAS participants are of European descent. Accordingly, most statistical methods developed to analyze GWAS data are designed under the implicit or explicit assumption that the population of interest is of a single homogeneous ancestry. Frequently, these methods perform suboptimally in applications that jointly consider multiple ancestry groups. This dissertation aims to demonstrate that there is great utility in the development of methods that take advantage of differences in across ancestries by jointly modeling diverse genetic datasets. In Chapter 1, we develop MultiSuSiE, a new statistical method for the identification of disease-causal variants that considers multiple ancestry groups simultaneously. Our method extends SuSiE, an existing single-ancestry method for causal variant identification. MultiSuSiE inherits the low computational cost of SuSiE, while increasing its power substantially by taking advantage of differences across diverse populations in the correlation structure of mutations. We apply MultiSuSiE to African and European ancestry whole-genome sequencing data from the All of Us dataset via simulations and analyses of real quantitative traits. Our method outperforms other methods in terms of power and computational expense. In Chapter 2, we integrate Latino ancestry data along with African and European ancestry data into our application of MultiSuSiE to the All of Us dataset. We demonstrate in simulations that MultiSuSiE is well-calibrated in the presence of possible sources of error not included in our Chapter 1 analyses and that the runtime of MultiSuSiE increases linearly with the number of ancestries. In real trait analyses, we use maximum sample sizes available to identify 82% more causal variants than in the Chapter 1 analyses. We further provide new recommendations for effective causal variation identification with Latino-ancestry data. In Chapter 3, we shed new light on a well-studied genetic phenomenon using diverse genetic data. Mutations which are more common in a population tend to have smaller effects on diseases and traits due to the action of negative selection. This phenomenon is typically modeled in a European dataset using the frequency of mutations in Europeans. We find two distinct lines of evidence which indicate that in Europeans, per-allele effect sizes are much better predicted using the frequency of mutations in African populations than European populations.