Publication: Generative Sequence Models for Variant Effect Prediction
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Interpreting the clinical significance of genetic variants remains a central challenge in genomics. While advances in sequencing have made it possible to identify millions of variants per individual, our ability to predict their functional and clinical consequences has not kept pace. Generative sequence models trained on evolutionary data spanning billions of years offer a principled approach to this problem by learning the constraints that maintain biological function without requiring labeled training data. This dissertation presents computational frameworks for variant effect prediction across three fundamental categories of genomic sequence: proteins, RNA, and regulatory DNA. For protein-coding variation, I present community guidelines for rigorous evaluation and dissemination of variant effect predictors, then introduce popEVE, a proteome-wide framework that integrates deep evolutionary information with human population data to produce calibrated pathogenicity scores comparable across all genes. Unlike existing methods that perform well within individual proteins but fail to generalize, popEVE enables identification of candidate disease variants directly from patient exomes, even in genes without prior disease associations. Application to developmental disorder cohorts identifies variants in 442 genes, including 123 novel candidates, many without requiring cohort-level statistical enrichment. For RNA, I present RNAGym, a comprehensive benchmarking framework integrating 70 deep mutational scanning assays with over one million variants, secondary structure data from chemical mapping experiments, and tertiary structures curated from the PDB. This addresses a critical gap in RNA variant interpretation where inconsistent evaluation practices have obscured comparative method performance across diverse RNA types, enabling systematic evaluation that reveals RNA secondary structure prediction substantially outperforms sequencebased fitness models. For regulatory variation, I present LOL-EVE, a conditional autoregressive transformer trained on mammalian promoter sequences that enables both insertion/deletion (indel) prediction and complete haplotype scoring. This addresses a critical gap in noncoding variant interpretation by demonstrating that evolutionary patterns learned from indels can accurately assess regulatory function, with application to clinical cohorts showing enrichment of deleterious promoter haplotypes in developmental disorder genes. Together, these contributions advance the goal of genome-wide variant interpretation by establishing rigorous evaluation standards, developing calibrated prediction frameworks, and extending evolutionary modeling approaches beyond coding sequences. The methods and benchmarks presented here provide foundations for potentially improving diagnostics in genetic disease and accelerating the translation of genomic information into clinical insight.