Skip to content

To ML-Predict or Not to ML-Predict: The Impact of Machine Learning-Predicted Protein Structures on FEP Accuracy and Data Augmentation.

Aug 2026 · Journal of Chemical Information and Modeling · Vol 66 17, pp. 10992-11005 · 0 citations · 36 references
Medicine

TL;DR

It is demonstrated that predictive variation in micro and macro conformational states─rather than the structural source─governs predictive reliability, underscoring the need for careful validation when integrating ML-derived structures into FEP workflows.

Abstract

The rapid advancement of machine learning (ML)-based protein structure prediction, exemplified by AlphaFold2 and extended by newer models such as AlphaFold3 and Boltz-2, has generated significant optimism for structure-guided drug discovery. In particular, ligand-protein cofolding approaches offer the potential to overcome limitations in generating starting structures for physics-based free energy perturbation (FEP) calculations. However, the practical readiness of ML-predicted structures for FEP applications remains insufficiently evaluated. Here, we systematically assess experimentally determined crystal structures, a homology model, and ML-predicted protein structures as inputs for FEP using a well-characterized congeneric series targeting the tyrosine kinase cSrc. A data set of 133 compounds was evaluated through more than 1400 FEP calculations under minimal optimization to approximate "out-of-the-box" performance. By maintaining consistent preparation protocols, we isolate the impact of structural origin on predictive accuracy. Variable performance was observed across both experimental and ML-predicted structures, highlighting that even under this idealized benchmark scenario, significant challenges remain in reliably generating and refining predictive protein-ligand complexes. This study demonstrates that predictive variation in micro and macro conformational states─rather than the structural source─governs predictive reliability, underscoring the need for careful validation when integrating ML-derived structures into FEP workflows.

View source

Similar papers

Sep 2026

Systematic benchmarking of AlphaFold and SWISS-MODEL kinase structures for structure-based drug discovery.

Accurate protein structure prediction is fundamental to structure-based drug discovery. However, the practical performance of deep learning-based models such as AlphaFold compared with traditional homology modeling approaches like SWISS-MODEL remains incompletely evaluated in realistic drug-design workflows. Here, we s...

Erick Bahena-Culhuac, Martiniano Bello · 0 citations
#protein folding Open access Sep 2026

De novo design of ligand binding proteins using large language models alone

Testing the ability of common large language models to consider design principles to generate de novo proteins that bind metals and lipophilic small molecules without copying existing sequences highlights the utility of LLMs in making protein design more comprehensible and accessible to users without sophisticated desi...

Nam Hyeong Kim, A. K. Hatstat, Hyunil Jo et al. · 0 citations
Open access Sep 2026

Comprehensive evaluation of AlphaFold/OpenFold prediction of experimentally unresolved proteins through novel metrics

Abstract Predicting accurate protein structures is essential for understanding molecular mechanisms, interpreting the impact of sequence variation, and supporting translational applications ranging from drug discovery to clinical genomics. Recent advances in deep-learning–based predictors such as AlphaFold2, OpenFold,...

Florencia R. Díaz, Daniela Orschanski, Juan I. Folco et al. · 0 citations
Jul 2026

Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2

Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.

Pascal Sturmfels, Naozumi Hiranuma, Milad Salem et al. · 0 citations
Open access Aug 2026

Boltz-Perturb: Improving Diversity and Accuracy in Protein-Ligand Co-Folding through Training-Free Conditioning Perturbation

Boltz-Perturb is presented, a framework for addressing small molecule binding poses through perturbing model conditioning signals during model inference, and it is demonstrated that inference-time perturbations can unlock latent structural diversity in generative co-folding models and improve protein-ligand predictions...

Hyeyun Jung, BoRam Lee, Alan C. Cheng · 1 citation
Aug 2026

Δ -Machine Learning for the Prediction of Metal Complex Properties.

An adaptation of the Δ-ML strategy for quantum property prediction of transition metal complexes is presented, which consistently achieves higher accuracy in the prediction of high-fidelity targets, while demonstrating improved data efficiency and out-of-domain transferability.

Hannes Kneiding, David Balcells · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.