Skip to content
Open access

A unified predictor of protein stability changes across all mutation types via implicit structure learning

Aug 2026 · Chemical Science · 0 citations · 58 references
Medicine

TL;DR

UniStab is introduced, an end-to-end framework for predicting stability changes across all mutation types by leveraging the implicit geometric reasoning of a pre-trained folding model and demonstrates state-of-the-art performance, particularly in the challenging scenarios of multi-point mutations and indels.

Abstract

Prediction of protein stability change caused by amino acid substitutions or indels (insertions/deletions) is crucial for protein engineering. While current models excel at single-point substitutions, they struggle with multi-point mutations and indels due to simplistic additivity assumptions and the inability to model backbone conformational changes. To address these limitations, we introduce UniStab, an end-to-end framework for predicting stability changes across all mutation types. By leveraging the implicit geometric reasoning of a pre-trained folding model, UniStab effectively captures non-additive epistatic interactions and local backbone rearrangements without the prohibitive cost of explicit structure generation. Evaluated on a comprehensive benchmark, UniStab demonstrates state-of-the-art performance, particularly in the challenging scenarios of multi-point mutations and indels. Beyond predictive accuracy, UniStab provides interpretable structural insights and effectively guides the design of stabilized variants, facilitating its potential utility in rational protein engineering.

Read PDF

Similar papers

Open access Aug 2026

An ensemble learning framework for protein stability prediction with enhanced recognition of stabilizing mutations

Three modeling frameworks are developed, including models based on handcrafted features, models using embedding representations extracted from ProteinMPNN, and ensemble models integrating a diverse set of state‐of‐the‐art predictors integrating a diverse set of state‐of‐the‐art predictors.

Yang Liu, Jian Zhang, Minghui Li · 0 citations
Open access Aug 2026

MAXWELL: Calibrating the probabilistic outputs of protein language models to the mutation-induced stability change landscape

Designing mutations that enhance protein stability is a central goal in protein engineering. However, experimentally screening large numbers of candidate mutations is costly and time-consuming, creating a strong need for computational methods that can identify potentially stabilizing mutations. Among these approaches, protein language models are particularly promising because they learn context-dependent amino acid preferences from large-scale sequence and structure datasets. Nevertheless, most existing stability prediction methods use these models primarily as feature extractors and do not fully exploit the amino acid probability distributions they encode. Here, we introduce MAXWELL (Matrix-wise Landscape Learning), a novel post-training method that calibrates the probabilistic outputs learned by protein language models during pretraining to generate mutational landscapes that quantify the effects of individual amino acid substitutions on protein stability. When applied to ProteinMPNN, MAXWELL yields a state-of-the-art predictor of the effects of protein mutations on stability, outperforming ThermoMPNN and other representative methods on a curated benchmark of experimentally measured stability changes. We next applied MAXWELL to the design of ten single-point mutations in the DhaA dehalogenase, seven of which (70%) increased thermal stability. Among them, G171W showed the largest improvement, with a measured ΔTm of 4.91 °C. These experimental results establish MAXWELL as a novel post-training strategy for protein language models and a practical framework for designing stabilizing mutations. Repository https://github.com/ai4protein/Venus-MAXWELL

Mingchen Li, Xiaoran Cheng, Fan Jiang et al. · 0 citations
Open access Aug 2026

Benchmarking Deep Learning Predictions of Mutation-Induced Fold Switching

A systematic NMR-characterized dataset of mutants of the GA/GB model fold-switching system is presented and it is found that this benchmark revealed variable and position-dependent performance across methods, with certain AlphaFold2-based algorithms able to predict mutant effects at individual sites, indicating some understanding of physical effects of residue substitutions.

Nathaniel R. Felbinger, K. Carillo, Yihong Chen et al. · 0 citations
Open access Jul 2026

Accurate ΔTm Prediction Without Protein Structure Inputs for Biomolecular Stability

Predicting protein stability, like changes in melting temperature (ΔTm) caused by mutations, is a critical task in therapeutic protein engineering and drug discovery. This is reflected by a growing solution space, including both AI-based sequence and structure based methods. This paper demonstrates that accurate ΔTm prediction does not require structural input features, but can achieve state-of-the-art results with a careful training design for large sequence-based protein language models. We combine an autoresearch-inspired setup search with controlled ablation studies and show that a well-tuned sequence-only ESM2-650M model [6] outperforms structure-informed methods in our benchmark, achieving the lowest error (MAE/RMSE) and competitive Pearson correlation without pH or structural inputs. We further show that choices such as loss function, pooling strategy, auxiliary supervision, and finetuning regime materially affect performance.

Daniel Siegismund, Mario Wieser, E. Natali et al. · 0 citations
Open access Aug 2026

Prediction of Distal Mutation Effects in Enzymes via Integration of Molecular Dynamics Descriptors and Zero-Shot Model

Identifying beneficial distal mutations remains a key challenge in enzyme engineering, as such residues can regulate activity through long-range dynamical coupling and allosteric communication. Although protein language models effectively capture sequence and evolutionary information, they lack explicit representation of conformational dynamics in the presence of the substrate, limiting their ability to detect distal regulatory sites. Here, we present an integrated framework that combines molecular dynamics (MD)-derived descriptors with the zero-shot prediction model GEMS to identify beneficial distal mutations. Compared to the pure zero-shot model, MD-derived descriptors efficiently capture key distal residues involved in dynamical coupling and allosteric communication with the active site. Consequently, these constraints enable the zero-shot model to predict distal mutations more precisely. By integrating sequence, evolutionary, and dynamic information, our approach expands the diversity of candidate sites while maintaining a manageable screening scale, offering an efficient and generalizable strategy for enzyme engineering.

Yiqiu Wang, Ding Luo, Shuming Cheng et al. · 0 citations