Accurately predicting how mutations alter protein–peptide binding remains challenging for molecular simulations because both conformational sampling and binding kinetics are computationally demanding. Here, we investigate whether multi-eGO, a hybrid transferable/structure-based atomistic-resolution model previously developed and validated for protein–small molecule interactions, can be transferred to protein–peptide binding without peptide-specific retraining. Using the PDZ2 domain of protein tyrosine phosphatase basophil-like in complex with the peptide EQVTAV as a benchmark, we first show that multi-eGO reproduces the structural dynamics of PDZ2 and the equilibrium binding thermodynamics of the wild-type complex. The model substantially accelerates both binding and unbinding relative to experiment but accurately preserves the resulting equilibrium dissociation constant. We then introduce conservative mutations in PDZ2 and in the peptide and evaluate their effects on binding without repeating the computationally expensive training procedure. Multi-eGO reproduces the experimentally observed changes in equilibrium dissociation constants, with strong agreement across PDZ2 mutants and moderate agreement when the peptide is also mutated. In contrast, the individual association and dissociation rate constants show substantially weaker agreement with experiment. The results indicate that the simplified energy landscape of multi-eGO limits the quantitative prediction of absolute kinetics while preserving thermodynamic information relevant to relative binding affinity. These findings establish multi-eGO as a computationally efficient approach for protein–peptide recognition and for predicting and rank-ordering the effects of conservative mutations on binding affinity.
Camilla Ardizzone, Bruno Stegani, Fran Bačić Toplek et al.· bioRxiv· 0 citations
The intrinsic dimension of a dataset is the number of independent directions needed to describe the space occupied by its data. Estimators based on nearest neighbors infer this number from how the probability to find a neighbor point grows around each sampled point. Because the distances $r$ between neighbor points decrease as the sample grows, the estimated dimension can depend strongly on the number of available points. Here, we derive the large-sample behavior of the TWO-NN estimator for data drawn from a smooth $d$-dimensional space. The typical nearest-neighbor distance scales as $N^{-1/d}$, and smooth deviations from a locally uniform distribution produce successive corrections proportional to $N^{-2/d}$. We test this result using the trajectories coming from ten independent $100~\mu$s simulations of alanine dipeptide. Configurations are represented by all pairwise distances among the ten heavy atoms. This representation has a known geometric dimension of $3n_{\mathrm{at}}-6=24$. Over the investigated range, the TWO-NN estimate shows no systematic dependence on the temporal spacing between configurations, but increases from approximately $7.5$ to $15.6$ as the sample size grows from $10^2$ to $2\times10^5$. Extrapolations that retain corrections through $r^2$, $r^4$, and $r^6$ give limiting dimensions of $25.23$, $22.89$, and $27.00$, respectively. All three estimates lie close to the known dimension and collectively bracket it, supporting the proposed scaling. Their spread provides a direct estimate of the systematic uncertainty associated with the truncation. The derived scaling therefore explains the strong sample-size dependence of TWO-NN and provides a practical route from finite sample estimates to the underlying geometric dimension.
Riccardo Capelli· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.