Publication:

Generative Sequence Models for Variant Effect Prediction

Loading...
Thumbnail Image

Date

2026-05-13

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Shearer, Courtney A. 2026. Generative Sequence Models for Variant Effect Prediction. Doctoral Dissertation, Harvard University Graduate School of Arts and Sciences.

Abstract

Interpreting the clinical significance of genetic variants remains a central challenge in genomics. While advances in sequencing have made it possible to identify millions of variants per individual, our ability to predict their functional and clinical consequences has not kept pace. Generative sequence models trained on evolutionary data spanning billions of years offer a principled approach to this problem by learning the constraints that maintain biological function without requiring labeled training data. This dissertation presents computational frameworks for variant effect prediction across three fundamental categories of genomic sequence: proteins, RNA, and regulatory DNA. For protein-coding variation, I present community guidelines for rigorous evaluation and dissemination of variant effect predictors, then introduce popEVE, a proteome-wide framework that integrates deep evolutionary information with human population data to produce calibrated pathogenicity scores comparable across all genes. Unlike existing methods that perform well within individual proteins but fail to generalize, popEVE enables identification of candidate disease variants directly from patient exomes, even in genes without prior disease associations. Application to developmental disorder cohorts identifies variants in 442 genes, including 123 novel candidates, many without requiring cohort-level statistical enrichment. For RNA, I present RNAGym, a comprehensive benchmarking framework integrating 70 deep mutational scanning assays with over one million variants, secondary structure data from chemical mapping experiments, and tertiary structures curated from the PDB. This addresses a critical gap in RNA variant interpretation where inconsistent evaluation practices have obscured comparative method performance across diverse RNA types, enabling systematic evaluation that reveals RNA secondary structure prediction substantially outperforms sequencebased fitness models. For regulatory variation, I present LOL-EVE, a conditional autoregressive transformer trained on mammalian promoter sequences that enables both insertion/deletion (indel) prediction and complete haplotype scoring. This addresses a critical gap in noncoding variant interpretation by demonstrating that evolutionary patterns learned from indels can accurately assess regulatory function, with application to clinical cohorts showing enrichment of deleterious promoter haplotypes in developmental disorder genes. Together, these contributions advance the goal of genome-wide variant interpretation by establishing rigorous evaluation standards, developing calibrated prediction frameworks, and extending evolutionary modeling approaches beyond coding sequences. The methods and benchmarks presented here provide foundations for potentially improving diagnostics in genetic disease and accelerating the translation of genomic information into clinical insight.

Description

Other Available Sources

Research Data

Keywords

AI, Bioinformatics, Generative Sequence Models, Human Disease, Synthetic Biology, Bioinformatics, Genetics, Artificial intelligence

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories