Jul 2026· Journal of Chemical Information and Modeling· Vol 66 15, pp.
9709-9720
· 0 citations· 19 references
Computer ScienceMedicine
TL;DR
EnerBridge-DPO, an energy-aware inverse folding framework that integrates Markov bridge sequence generation with preference optimization for protein complex design, suggests that incorporating energy-related preferences into Markov bridge inverse folding can improve computationally predicted energetic profiles.
Abstract
Designing protein sequences with favorable predicted energetic properties is an important challenge in protein inverse folding, because many existing deep learning methods are primarily trained by maximizing sequence recovery and do not explicitly incorporate energy-related preferences during generation. In this work, we propose EnerBridge-DPO, an energy-aware inverse folding framework that integrates Markov bridge sequence generation with preference optimization for protein complex design. The framework builds on the Markov bridge inverse-folding process to generate structure-compatible sequences from an informative prior sequence. It then introduces a Bridge-DPO objective that uses energy-related winner-loser preference pairs to bias the generator toward sequences favored by computational or experimental energy-related signals. In addition, we incorporate a quantitative energy-constrained loss based on mutation-induced binding free-energy changes to provide continuous ΔΔG supervision. Evaluations show that EnerBridge-DPO maintains competitive inverse-folding performance while obtaining lower predicted energy scores under selected computational scoring functions for protein complexes. On SKEMPI, EnerBridge-DPO achieves competitive ΔΔG prediction performance, with small numerical gains in several overall metrics that are not statistically conclusive under paired bootstrap analysis. These results suggest that incorporating energy-related preferences into Markov bridge inverse folding can improve computationally predicted energetic profiles, although experimental validation is required to confirm thermodynamic stability.
This work introduces a symmetric dual-path architecture that both leverages PLMs for pretrained sequence evolution knowledge and MPLMs for pretrained structural knowledge to iteratively guide protein sequence generation.
Han-Dong Wang, Jiaxin Qi, Baisheng Lai et al.· 0 citations
TTS-Design is proposed, a test-time compute scaling framework that enhances protein sequence design without retraining models or relying on larger training datasets, and can consistently improve sequence recovery and structural reliability across different backbone models, without retraining or increasing model size.
Zizhe Jin, Yi Zheng, Huan Yee Koh et al.· Proceedings of the Thirty-Fi...· 0 citations
A function-aware preference alignment framework that improves functional preservation by fine-tuning models to favor function-preserving sequences over function-disrupting alternatives, avoiding the need for explicit function optimization.
Nilufer Tamatgar, Soobin Park, Ying-Hua Yao et al.· Proceedings of the 32nd ACM...· 0 citations
Inverse FoldDir is a structure-conditioned protein redesign method that combines structural recovery, user control, experimental validation, and a natural route toward future property-guided sampling that performs iterative denoising on the amino acid probability simplex.
Alp Tartici, M. Stojkovic, An-Ru Tian et al.· bioRxiv· 0 citations
It is shown that blending IF models with a physics-based coarse-grained potential improves global correlation with experimental ΔΔG and, crucially, reduces IF model bias at functional sites, and is found that disease gain-of-function variants show a distinct functional signature from loss-of-function variants.
Ezequiel A. Galpern, Xavier Soler Sanchis, Charles W. J. Pugh et al.· bioRxiv· 0 citations
Predicting the impact of single-point mutations on protein thermodynamic stability is crucial for protein engineering of therapeutic and industrial applications. By effectively capturing the three-dimensional structural information of proteins and the spatial physical environment of each residue, the inverse folding models (IFMs) upon fine-tuning, such as ThermoMPNN, achieved state-of-the-art performance in predicting thermostability changes in proteins caused by mutations. However, IFMs are limited in their capacity to capture protein deep evolutionary information, whereas protein language models (pLMs) excel. Here, we present SA-MPNN, a lightweight, end-to-end hybrid framework that dynamically integrates the protein sequence representations from a protein language model (ESM2) into the ThermoMPNN architecture to improve protein stability prediction by combining evolutionary representations with geometric structural embeddings. By evaluating various feature fusion strategies, we selected a self-attention-based integration mechanism to effectively combine the two modalities. Trained on the large-scale Megascale data set, SA-MPNN achieved modest but consistent gains over ThermoMPNN on various benchmark data sets, with particularly noticeable improvements in several correlation analysis and screening-oriented evaluations. Finally, wet-lab validation was performed on the top-ranking variants of Acetivibrio thermocellusβ-glucosidase (AtBgl1A) as a case study. The experimental results demonstrated that multiple designed mutants exhibited improved thermostability, and the optimal variant, GC20, achieved a melting temperature (Tm) of 76.98 °C, representing a 5.97 °C increase over the wild-type, thereby supporting the practical applicability of SA-MPNN in protein engineering.
Xin-Yue Zhang, Xiangshan Zheng, Ze-Yuan Dong et al.· Journal of Chemical Informat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.