Skip to content
Preprint

Navigating committor landscape of biomolecules with a general pairwise interaction model

Jun 2026 · 0 citations · 4 references
Physics

TL;DR

A novel committor learning framework grounded in the AlphaFold 3 paradigm is proposed that elucidates how ligand substituents regulate the ratio between distinct binding pathways, offering new perspectives for structure-based drug design.

Abstract

Sampling rare conformation transitions between metastable states is a central challenge in atomistic simulations. While the committor function serve as an ideal reaction coordinate for driving enhanced sampling, their high-dimensional inputs and complex functional forms limit the efficacy of standard feedforward neural networks in modeling them. Inspired by recent breakthroughs in biomolecular structure prediction, we propose a novel committor learning framework grounded in the AlphaFold 3 paradigm. By integrating a lightweight, differentiable atom-level embedding with a simplified Pairformer architecture, our method inherently captures intricate dynamical features of diverse biosystems without requiring specialized prior knowledge. We demonstrate the superior expressiveness and accuracy of the proposed framework across multiple atomistic processes. For the folding of the chignolin mini-protein, our model reveals the finer-grained structure of its transition state ensemble (TSE) and a detailed bifurcated reaction mechanism. Furthermore, for calixarene host-guest systems, we develop a unified committor model that elucidates how ligand substituents regulate the ratio between distinct binding pathways, offering new perspectives for structure-based drug design.

View source

Similar papers

Preprint Aug 2026

PHASE: encoding global protein ensembles with local Hamiltonians and all-atom backmapping

Protein function is governed by conformational ensembles, which can be viewed as high-dimensional probability distributions over molecular conformations. Yet the statistical organization of these distributions is often represented only implicitly, either through collections of simulation trajectories or within high-capacity generative models. Here, we introduce PHASE (Protein Hamiltonians for Sampling of Ensembles), a system-specific framework that converts atomistic conformational ensembles into an explicit and interpretable statistical model. Applied to ten conformational ensembles derived from approximately 37$\mu$s of atomistic simulations of the adenosine A2A receptor, Hamiltonians containing only local residue couplings within 6$\mathring{A}$ reproduce residue-wise and pairwise microstate statistics, including correlations between residues that are not directly coupled in the model. Moreover, independently fitted inactive and active reference Hamiltonians define an endpoint preference coordinate that organizes newly sampled ligand-, effector- and conformation-dependent ensembles along the A2A activation landscape without receiving these biochemical labels as model inputs. Finally, a cluster-conditioned all-atom reconstruction model preserves the prescribed residue microstate patterns of newly sampled configurations, closing the coarse-graining-sampling-backmapping cycle. The resulting discrete representation additionally admits direct QUBO encoding, enabling classical annealing and providing a route toward future quantum-annealing implementations. PHASE therefore provides a protein-general procedure for constructing compact, interpretable and atomistically realizable statistical models of protein conformational ensembles.

Daniele Angioletti, Marco S. Nobile, Matteo Carli et al. · 0 citations
Open access Aug 2026

Pi-Ensemble: Sequence-guided generation of interpolated protein conformational ensembles

Proteins are critical biomolecular machines that populate ensembles of interconverting conformations. Many biological processes depend on transitions between metastable states. Although molecular dynamics (MD) simulations provide a physically grounded route to characterize these motions, routine sampling of large-scale conformational transitions remains computationally demanding. Recent advances in protein structure prediction have created new opportunities for ensemble generation, but many existing approaches require noising inputs, task-specific training, supervised fitting on extensive MD data, or experimentally-informed restraints. Here, we introduce Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states. Unlike previous methods, Pi-Ensemble alternately leverages inverse-folding and structure-prediction models to propose intermediate conformations between known protein states, generating diverse ensembles without additional training. We evaluate Pi-Ensemble across diverse protein systems, including enzymes, transporters, receptors, and benchmark cases with reference MD simulations or experimental Double Electron-Electron Resonance (DEER) data. Pi-Ensemble recovers physically plausible intermediate conformations, captures transition pathways observed in large-scale MD simulations, and generates structures consistent with experimental distance distributions. Furthermore, Pi-Ensemble-generated conformations provide effective starting seeds for parallel MD simulations, improving conformational exploration and accelerating convergence relative to simulations initiated only from endpoint structures. These results establish sequence-guided structural interpolation as a practical strategy for probing protein conformational landscapes. By generating diverse and physically reasonable conformational proposals without long-timescale MD or model retraining, Pi-Ensemble provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.

Hassan Nadeem, D. Kleiman, Yuming Zhou et al. · 0 citations
Open access Aug 2026

Transferable Collective Variable to accelerate Protein-Ligand (Un)Binding Transitions via Explainable Machine Learning and Intriguing Role of Ligand Solvation

The process of drug unbinding is of immense importance in the field of biophysics and therapeutics. The behavior of these systems is greatly influenced by their thermodynamic and kinetic properties. Therefore, it is crucial to accurately estimate the ligand binding free energies and rate of ligand dissociation, yet these processes are often governed by rare event transitions that lie beyond the reach of standard brute-force molecular dynamics simulations. While enhanced sampling simulations offer a solution, their efficacy is strictly contingent upon the selection of appropriate collective variables (CVs) which is non-trivial for complex systems like protein-ligand complexes. In this study, we present a method to derive optimized CV from transition state region (TS) via an interpretable machine learning (ML) model, Elastic Net. By employing some physically intuitive order parameters, the derived optimized CV from the TS-region greatly accelerate ligand binding-unbinding transitions and achieves rapid free energy surface (FES) convergence across diverse systems including buried and solvent exposed active sites such as Trpsin-benzamidine complex, host-guest systems and sodium epoxidase etc. Intriguingly significant contribution of the ligand hydration is found in the optimized CV which depicts crucial role of solvent in driving ligand binding-unbinding transitions. The estimated binding free energies for different protein-ligand complexes match quite well with experiments, while maintaining a low computational cost. The derived optimized CV is also used to calculate the ligand residence times across different systems and calculated residence times are within the experimental range for all systems, again with very little computational costs. Moreover, we show that the optimized CV constructed from TS region via an interpretable ML model is transferable across diverse systems, offering a robust and scalable framework for drug discovery and investigation of complex biomolecular recognition.

Saikat Dhibar, B. Jana · 0 citations
Open access

Experimental data-guided parameterization and validation of an AMBER protein force field

Folded or globular proteins adopt well-defined three-dimensional structures that correlate directly with function, forming the classical structure-function paradigm. Intrinsically Disordered Proteins (IDPs) and Regions (IDRs), encoded by roughly one-third of the eukaryotic genome, lack this fixed structure; their structural plasticity and major roles in diverse biological phenomena instead challenge this paradigm, making them a critical class of biomolecules studied through both experimental and computational approaches. Molecular dynamics (MD) simulations provide an atomistic understanding of IDP structure and dynamics, but their accuracy is limited by methodological assumptions and by the precision of the force field--the potential-energy function that approximates atomic interactions and governs how faithfully a simulation reflects real conformational behavior. Previous studies show that contemporary force fields fail to adequately capture amino acid residue specificity, limiting their applicability to IDPs. The central focus of this dissertation is therefore the development of an improved force field, Amber ff24EXP-GA, derived from its parent, Amber ff14SB, and its evaluation against Amber ff14SB and other contemporary force fields, such as CHARMM36m, in capturing the empirically determined conformational properties of unfolded systems: short peptides that serve as model systems for IDPs, and longer unfolded proteins. Chapter 3 addresses this gap by reporting a new force field, Amber ff24EXP-GA, derived from Amber ff14SB by optimizing backbone dihedral potentials for guest glycine and alanine residues in cationic GGG and GAG peptides, respectively, to best match guest-residue-specific spectroscopic data. Amber ff24EXP-GA outperforms Amber ff14SB for conformational ensembles of all 14 guest residues x (G, A, L, V, I, F, Y, D^p, E^p, R, C, N, S, T) in GxG peptides in water--the full set for which spectroscopic data exist--and outperforms CHARMM36m for at least 7 of them (G, A, V, F, C, T, E^p), showing greater amino acid specificity than both Amber ff14SB and CHARMM36m. It also reproduces experimental data on three-folded proteins and three longer IDPs well, while still outperforming Amber ff14SB on short unfolded peptides. Chapter 4 examines the effect of nearest-neighbor (NN) residues on conformational dynamics: extensive experimental evidence shows that the Flory hypothesis--which assumes neighboring residues behave independently--does not hold, as NN residues instead alter the conformational landscape of a given residue. Here, we evaluate CHARMM36m, Amber ff14SB, and Amber ff24EXP-GA for their ability to capture these NN effects on conformational dynamics of amino acid residues in short unfolded peptides in water. Amber ff24EXP-GA, whose reproduction of intrinsic conformational ensembles is significantly better than its parent's, also captures the NN effects better than Amber ff14SB. Despite lacking residue specificity in its intrinsic conformational ensembles, CHARMM36m--calibrated on global IDP properties--reproduces the NN effects on par with Amber ff24EXP-GA. These findings matter for the development of next-generation force fields that capture both residue-specific dynamics and global IDP properties. Chapter 5 examines the current scope, achievements, and limitations of various MD force-field parametrization strategies for modeling proteins, with particular emphasis on intrinsically disordered proteins.

Athul Suresh, B. Urbanc · 0 citations
Aug 2026

Integrating Parallel-Bias Metadynamics-Metainference and Relaxation Dispersion NMR to Resolve Hidden Conformational States Driving Toxin-Channel Recognition.

Conformational dynamics in toxin inhibitors are an important contributor to ion channel affinity, yet toxin multidimensional energy landscapes remain largely unexplored. In the current work, we combine parallel-bias metadynamics-metainference (PBMetaD) simulations with relaxation dispersion NMR to define, at atomistic resolution, the thermodynamics and kinetics of Hui1, a de novo three disulfide toxin derived from the SAK-I family that targets K+-channels. Using the three χ3 disulfide dihedrals as collective variables, an extensive 48-replica well-tempered PBmetaD simulation (16.2 μs cumulative sampling) resulted in a fully converged three-dimensional (3D) free-energy surface comprising eight Hui1 conformers. These basins account for ∼96% of the bias-weighted ensemble and partition into four low- and four high-energy states separated by 7.5 kJ/mol associated with the (-) and (+)Cys12-Cys28 χ3 states, respectively. Transition-state theory identifies rota-isomerization of Cys3-Cys35 as the slowest, and therefore rate-determining, coordinate, while the analogous motions around Cys12-Cys28 and Cys17-Cys32 are ∼5-fold faster. 15N R1ρ relaxation dispersion NMR measurements confirmed these kinetics, identifying two structurally distinct residue clusters exhibiting intermediate and faster exchange processes. The Key Interaction Finder (KIF) approach reveals that Cys17-Cys32 conformation modifies connectivities within a dense interaction network between Cys17 and residues Gln14, Tyr23, Arg24, and Lys29, and correlates with the accessibility of residues Tyr23 and Arg24 of the helix-kink-helix region for interaction with the channel vestibule. Our work establishes PBMetaD as a powerful framework for mapping coupled disulfide and backbone dynamics in toxins and reveals specific conformers and interactions likely to control K+ channel recognition.

Chen Timsit Shmueli, Miriam Gulman, D. T. Major et al. · 0 citations