Publication:

Decoding the pathogenicity of epilepsy-associated non-coding genetic variants

Loading...
Thumbnail Image

Date

2026-05-19

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Zhang, Lanxin. 2026. Decoding the pathogenicity of epilepsy-associated non-coding genetic variants. Masters Thesis, Harvard Medical School.

Abstract

Epilepsy is among the most prevalent neurological disorders, affecting an estimated 50 million people worldwide (World Health Organization, 2024), and is increasingly recognised to have a substantial genetic component (Thomas & Berkovic, 2014; Epi25 Collaborative, 2024). A central interpretive challenge in epilepsy genetics is the prioritization of non-coding variants, which constitute most variation discovered by whole-genome sequencing yet rarely have an established functional or clinical interpretation (Maurano et al., 2012; Ward & Kellis, 2012). Variants in regulatory elements such as enhancers, promoters, and splice-affecting non-coding regions can disrupt gene expression and contribute to disease, but their effects are not predictable from sequence alone with commonly-used methodologies. In our study, we present a scalable, multi-modal computational approach, termed BREAD (Biological Regulatory Element Aggregation for variant Discovery), to prioritize potentially pathogenic non-coding variants. We applied this pipeline to two epilepsy-associated genes, SYNGAP1 and SLC12A5, and managed to filter from a total of 738 and 1,071 non-coding variants to 49 and 27 candidate variants for further testing, respectively. Our approach uses population-scale variant data, neural epigenomic annotations, and deep learning-based regulatory element predictions. Our pipeline reports that splicing disruption is the most common regulatory mechanism, with 76 and 108 high-impact variants in SYNGAP1 and SLC12A5, respectively, associated with splicing effects. Our study also demonstrates transcription factor binding site disruption, and indeed, 46 and 9 variants in SYNGAP1 and SLC12A5, respectively, were associated with significant motif disruption. When combined with evolutionary constraint, two high-confidence variants in SYNGAP1, namely, chr6:33432030:T>TAGGTGAGGC and chr6:33448856:C>T, were associated with splicing disruption, transcription factor binding site disruption, and strong phyloP (> 3) constraint. In summary, the BREAD analysis pipeline is highly adaptable and scalable, enabling mechanistic insights that drive genetic variant discovery and prioritization in clinical settings. It has the potential to help realizing the promise of genomic medicine by transforming the growing abundance of genomic data into actionable insights.

Description

Other Available Sources

Research Data

Keywords

Bioinformatics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories