KinConfBench is introduced, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states.
Abstract
Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.
Boltz is benchmarked using a curated set of ligand-bound human G protein-coupled receptors from families unseen during training, showing that while Boltz generally predicts receptor backbones accurately, ligand poses can contain significant errors that lead to a limited ability to reproduce experimental affinity data w...
Lichirui Zhang, R. Friesner, Edward B. Miller et al.· npj Drug Discovery· 0 citations
A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is pres...
Ryan Varghese, Pooja Tiwary, Krishil Oswal· bioRxiv· 0 citations
This work systematically benchmarked several molecular docking tools representing all-atom models, highlighting the complementarity of AI-based approaches and methods based on physical sampling in all-atom models, in terms of their applicability range, and argues for the benefits of tighter integration.
Ao Xu, J. Lam, A. Nakano et al.· Journal of Chemical Informat...· 0 citations
Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.
Pascal Sturmfels, Naozumi Hiranuma, Milad Salem et al.· arXiv.org· 0 citations
It is shown that prospective structure selection, rather than structure generation, represents the primary bottleneck in ensemble-based VS, highlighting an urgent need for novel structural descriptors to identify high-performing conformations.
Jaeoh Shin, K. Joo, Jejoong Yoo· Journal of Chemical Informat...· 1 citation
It is demonstrated that BioEmu can generate plausible conformational ensembles for relatively large, six-and seven-pass membrane proteins, sampling rare states at a fraction of the computational cost of conventional MD simulations, suggesting that AI-based ensemble generation could provide an accessible approach for ex...
B. Clifton, Adam G. Grieve, Robin A. Corey· bioRxiv· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.