Publication:

Evaluating the Impact of Training Set Composition on Protein Structure Prediction Model Sensitivity to Mutations

Loading...
Thumbnail Image

Date

2026-06-02

Published Version

Published Version

Journal Title

Journal ISSN

Volume Title

Publisher

The Harvard community has made this article openly available. Please share how this access benefits you.

Research Projects

Organizational Units

Journal Issue

Citation

Egbebi, Oluwatoyosi. 2026. Evaluating the Impact of Training Set Composition on Protein Structure Prediction Model Sensitivity to Mutations. Bachelors Thesis, Harvard University Engineering and Applied Sciences.

Abstract

Machine-learning-based structure prediction models achieve high accuracy on general sets of wild-type proteins, but many exhibit limited sensitivity to representing mutations in their predicted structures. This study investigates whether modifying the proportion of a target protein in training data can influence mutation sensitivity in structure prediction models, using the serum albumin protein family as a case study. Two versions of the Boltz-1 structure prediction model were partially retrained: a baseline model reflecting "default" training set composition, and an albumin-enriched model with ~25% increased albumin representation. Model performance was evaluated using curated validation (wild-type) and test (mutant) sets, using metrics including RMSD and TM score to quantify the similarity of predicted structures to the wild-type proteins. Results showed that the pretrained Boltz-1 model was the highest performing overall, while the partially retrained models exhibited more variability in predictions. Despite their lower performance, the retrained models showed differential performance in some cases, like predicting the structures of wild-type and variants of human serum albumin. These findings suggest that training dataset composition can influence model behavior, although continued training would ideally improve prediction accuracy and consistency. Overall, this study highlights the considerations and challenges of retraining large-scale structure prediction models for the purpose of evaluating mutation sensitivity.

Description

Other Available Sources

Research Data

Keywords

Boltz-1, machine learning, mutations, protein engineering, structure prediction, Biomedical engineering, Computer science

Terms of Use

This article is made available under the terms and conditions applicable to Other Posted Material (LAA), as set forth at Terms of Service

Endorsement

Review

Supplemented By

Related Stories