Intrinsically disordered proteins (IDPs) drive diverse cellular processes through broad conformational ensembles, but experimental characterization of these ensembles is sparse and high-quality training data is scarce, holding back artificial intelligence (AI) approaches to ensemble prediction. Single-bead coarse-grained (CG) force fields such as CALVADOS, parametrized directly against experimental observables, currently provide a more reliable route to disordered ensembles than direct AI prediction. CG sampling lacks atomistic resolution and must be paired with backmapping; the atomistic information recoverable from this two-step process is shaped jointly by the CG representation and the backmapping algorithm, and the interplay between these contributions is not well characterized. We benchmarked CODLAD, a recently published latent-diffusion backmapping pipeline, on 23 Protein Ensemble Database (PED) systems using a dual-input design: the same architecture receives either PED-reference Cα coordinates or independently sampled CALVADOS Cα trajectories. This design separates paired reconstruction error, measurable for PED+CODLAD, from unpaired CG-input-associated ensemble deviations, measurable for CG+CODLAD at the distribution level. Reconstruction from PED conformers achieved 0.56 ± 0.08 Å backbone root-mean-square deviation and preserved local geometry. Reconstruction from CALVADOS-sampled Cα trajectories preserved the Cα framework and bond geometry, while ensemble-level deviations relative to PED were localized to proline backbone geometry and cooperative secondary structure, both consistent with information not carried by an unconstrained single-bead representation. The benchmark quantifies which atomistic properties can be recovered after projecting a CG ensemble into all-atom space and identifies improved CG geometric encoding and sequence-conditioned AI priors as the directions for further progress.
Jianxiang Huang, Xin Qiao, Ning Liu et al.· Journal of Chemical Informat...· 0 citations
Abstract Cyclic peptides represent a highly promising class of biopharmaceutical scaffolds. The screening of cyclic peptides against protein targets can be greatly facilitated using computational approaches, especially molecular docking. However, it remains a crucial challenge to accurately predict protein–cyclic peptide (P–cp) interactions employing scoring functions of molecular docking. End-point approaches, such as molecular mechanics generalized Born surface area (MM/GBSA) and molecular mechanics Poisson–Boltzmann surface area (MM/PBSA), provide theoretically more robust frameworks than conventional scoring functions, but their reliability in predicting binding affinities and discriminating native-like binding poses for P–cp complexes remains poorly quantified. Herein, we comprehensively assessed the predictive abilities of MM/PBSA(GBSA) in scoring binding affinities of P–cp complexes and re-ranking their binding poses. The binding affinity scoring ability of MM/PBSA(GBSA) was assessed on a carefully curated dataset consisting of 50 complexes involving P–cp binding affinities, and their re-ranking capability was evaluated on another dataset consisting of the decoys of 81 P–cp complexes. Based on these assessments, we proposed a two-step workflow for predicting P–cp binding affinities. First, we employed the assessed optimal re-ranking method to select the top-1 binding pose; second, we estimated the binding affinity based on the selected top-1 pose using the assessed optimal scoring method. Our proposed workflow, which requires only 3 s for each prediction, achieves binding affinity predictions with a Rp of −0.732 when compared to experimental values, which is twice as high as that of AutoDock CrankPep (Rp = −0.316). This study emphasizes the necessity of using fine-tuned MM/PBSA(GBSA) methods for predicting P–cp interactions.