Publication: Evaluating the Impact of Training Set Composition on Protein Structure Prediction Model Sensitivity to Mutations
Open/View Files
Date
Authors
Published Version
Published Version
Journal Title
Journal ISSN
Volume Title
Publisher
Citation
Abstract
Machine-learning-based structure prediction models achieve high accuracy on general sets of wild-type proteins, but many exhibit limited sensitivity to representing mutations in their predicted structures. This study investigates whether modifying the proportion of a target protein in training data can influence mutation sensitivity in structure prediction models, using the serum albumin protein family as a case study. Two versions of the Boltz-1 structure prediction model were partially retrained: a baseline model reflecting "default" training set composition, and an albumin-enriched model with ~25% increased albumin representation. Model performance was evaluated using curated validation (wild-type) and test (mutant) sets, using metrics including RMSD and TM score to quantify the similarity of predicted structures to the wild-type proteins. Results showed that the pretrained Boltz-1 model was the highest performing overall, while the partially retrained models exhibited more variability in predictions. Despite their lower performance, the retrained models showed differential performance in some cases, like predicting the structures of wild-type and variants of human serum albumin. These findings suggest that training dataset composition can influence model behavior, although continued training would ideally improve prediction accuracy and consistency. Overall, this study highlights the considerations and challenges of retraining large-scale structure prediction models for the purpose of evaluating mutation sensitivity.