Tang, XinNaqvi, SahinZhang, Lanxin2026-05-1920262026-05-192026Zhang, Lanxin. 2026. Decoding the pathogenicity of epilepsy-associated non-coding genetic variants. Masters Thesis, Harvard Medical School.32699867https://dash.harvard.edu/handle/1/42738362Epilepsy is among the most prevalent neurological disorders, affecting an estimated 50 million people worldwide (World Health Organization, 2024), and is increasingly recognised to have a substantial genetic component (Thomas & Berkovic, 2014; Epi25 Collaborative, 2024). A central interpretive challenge in epilepsy genetics is the prioritization of non-coding variants, which constitute most variation discovered by whole-genome sequencing yet rarely have an established functional or clinical interpretation (Maurano et al., 2012; Ward & Kellis, 2012). Variants in regulatory elements such as enhancers, promoters, and splice-affecting non-coding regions can disrupt gene expression and contribute to disease, but their effects are not predictable from sequence alone with commonly-used methodologies. In our study, we present a scalable, multi-modal computational approach, termed BREAD (Biological Regulatory Element Aggregation for variant Discovery), to prioritize potentially pathogenic non-coding variants. We applied this pipeline to two epilepsy-associated genes, SYNGAP1 and SLC12A5, and managed to filter from a total of 738 and 1,071 non-coding variants to 49 and 27 candidate variants for further testing, respectively. Our approach uses population-scale variant data, neural epigenomic annotations, and deep learning-based regulatory element predictions. Our pipeline reports that splicing disruption is the most common regulatory mechanism, with 76 and 108 high-impact variants in SYNGAP1 and SLC12A5, respectively, associated with splicing effects. Our study also demonstrates transcription factor binding site disruption, and indeed, 46 and 9 variants in SYNGAP1 and SLC12A5, respectively, were associated with significant motif disruption. When combined with evolutionary constraint, two high-confidence variants in SYNGAP1, namely, chr6:33432030:T>TAGGTGAGGC and chr6:33448856:C>T, were associated with splicing disruption, transcription factor binding site disruption, and strong phyloP (> 3) constraint. In summary, the BREAD analysis pipeline is highly adaptable and scalable, enabling mechanistic insights that drive genetic variant discovery and prioritization in clinical settings. It has the potential to help realizing the promise of genomic medicine by transforming the growing abundance of genomic data into actionable insights.application/pdfenBioinformaticsDecoding the pathogenicity of epilepsy-associated non-coding genetic variantsThesis or Dissertation2026-05-190009-0007-1169-9418