Publication:

Statistical methods for polygenic risk prediction from biobanks and genome-wide association studies

Loading...
Thumbnail Image

Date

2025-06-05

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Chen, Tony Antong. 2025. Statistical methods for polygenic risk prediction from biobanks and genome-wide association studies. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Polygenic risk scores (PRS) have emerged as a promising tool to translate genomic discoveries into clinic decision making. By aggregating the effects of risk-associated genetic variants across the genome into a single number, PRS can quantify a patient’s genetic predisposition for a wide range of health outcomes. The past decade has seen an explosion of statistical methodology to build PRS prediction models, but these existing methods face limitations in prediction accuracy, computational efficiency, and generalizability across populations. In this dissertation, we present three novel approaches to tackle these challenges.

In Chapter 1, we propose ALL-Sum, an ensemble learning-based PRS method that uses summary statistics from genome-wide association studies (GWAS). ALL-Sum leverages L0L2-penalized regression and fast optimization algorithms to enable high prediction accuracy while also dramatically reducing the computational runtime and memory usage. Then, in Chapter 2, we propose SPLENDID, which models gene-by-ancestry interactions to simultaneously capture shared and heterogeneous genetic effects without categorizing individuals into discrete ancestry groups, allowing for fairer clinical implementation in diverse patient populations. Finally, in Chapter 3, we propose STELLAR, which ensembles multiple prediction modeling approaches and functional genomic annotations to flexibly estimate rare variant effects for more comprehensive genome-wide PRS.

Altogether, the work from this dissertation introduces new statistical frameworks to efficiently compute accurate PRS from large-scale genomic data. Each method demonstrates substantial improvements over the current state-of-the-art through comprehensive simulation studies and application to real data. These tools can be used to develop new prediction models for complex traits and diseases and ultimately advance precision medicine.

Description

Other Available Sources

Research Data

Keywords

Biostatistics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories