Publication:

Compositional Phenotype Representations and Clinical Prior Distillation for Missense Variant-Disease Association Retrieval

Loading...
Thumbnail Image

Open/View Files

Date

2026-05-15

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Cai, Shuo. 2026. Compositional Phenotype Representations and Clinical Prior Distillation for Missense Variant-Disease Association Retrieval. Masters Thesis, Harvard Medical School.

Abstract

Missense variants are by far the most prevalent type of pathogenic coding variants, but they are predominantly annotated as variants of uncertain significance (VUS) in ClinVar. This categorization prevents any medical intervention despite the variant being causal. Ex- isting computational tools predict pathogenicity well but cannot answer the question that drives genetic diagnosis: harmful for which disease? We present PheMART2, a contrastive retrieval framework that ranks candidate diseases for each query missense variant. To solve these problems, we implement two major architectural designs to improve general- ization. Diseases are encoded using weighted attention-based combinations of HPO terms instead of relying on concept IDs, allowing the model to generalize to disease labels outside the training distribution. Population-scale variant–phenotype relationships from the MVP phenome-wide association study are mapped to the HPO disease space through a curated CUI mapping and SapBERT-based semantic translation pipeline, and then distilled into our model using bidirectional Kullback-Leibler divergence in both the variant→disease and disease→variant directions. Model evaluation is performed by randomly splitting data on the basis of genes instead of per-variant random splits with negative sampling, across all 1,563 diseases in the dataset. In comparison with this more stringent evaluation approach, PheMART2 demonstrates an MRR of 0.272 for unseen genes, an 18.1% improvement over the no-KL, no-auxiliary version (p = 0.001), with the clinical prior distillation being the key contribution. Disease label out-of-distribution inference shows an MRR score of 0.39, suggesting that our compositional encoding successfully generalizes to diseases constructed from novel phenotypic primitives. Amongst 676 independently curated Phenopackets cases with no overlap between variants used in the training data, the top 10 candidate list in- cludes the true disease for 51.5% and the top 1 for 18.5%. Performance concentrates on variants in genes with existing biological annotations (MRR = 0.418) and falls sharply on genes without much connection in the knowledge graph (MRR = 0.089), which reflects the inherent constraint of the transductive graph setting and points to inductive extension as the natural next step.

Description

Other Available Sources

Research Data

Keywords

Bioinformatics

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories