The biosynthetic machineries of ribosomally synthesized and post-translationally modified peptides (RiPPs) are often substrate tolerant. A remarkable example is the class II lanthipeptide synthetase ProcM, which naturally functions as a generalist enzyme that has not evolved to use a specific substrate during its evolutionary history. Although ProcM has been studied extensively, the sequence features associated with productive modification remain underexplored. In this study, we use the ultrahigh-throughput mRNA display technique to map the sequence compatibility of ProcM across a focused library. This approach expands the landscape of ProcM reactivity beyond native substrates and individually characterized variants. Machine learning (ML) is used as a tool to demonstrate that the selected dataset contains learnable signatures and classification architectures revealed a balanced accuracy of 0.73. This performance contrasts sharply with the near-perfect accuracy of specialized enzyme models as the sequence-fitness landscape of the generalist enzymes are characterized by class imbalance and limited by intrinsic dataset features. Our results provide a high-throughput view of ProcM reactivity and highlight differences with previous high-throughput studies on substrate selectivity of RiPP modification enzymes. Future studies will need to assess whether these differences are common when comparing generalist with specialist enzymes. Table of Contents Graphic
Proteins are critical biomolecular machines that populate ensembles of interconverting conformations. Many biological processes depend on transitions between metastable states. Although molecular dynamics (MD) simulations provide a physically grounded route to characterize these motions, routine sampling of large-scale conformational transitions remains computationally demanding. Recent advances in protein structure prediction have created new opportunities for ensemble generation, but many existing approaches require noising inputs, task-specific training, supervised fitting on extensive MD data, or experimentally-informed restraints. Here, we introduce Pi-Ensemble (Predicting Interpolated Ensemble), a sequence-guided framework for generating protein conformational ensembles interpolating between two structural anchor states. Unlike previous methods, Pi-Ensemble alternately leverages inverse-folding and structure-prediction models to propose intermediate conformations between known protein states, generating diverse ensembles without additional training. We evaluate Pi-Ensemble across diverse protein systems, including enzymes, transporters, receptors, and benchmark cases with reference MD simulations or experimental Double Electron-Electron Resonance (DEER) data. Pi-Ensemble recovers physically plausible intermediate conformations, captures transition pathways observed in large-scale MD simulations, and generates structures consistent with experimental distance distributions. Furthermore, Pi-Ensemble-generated conformations provide effective starting seeds for parallel MD simulations, improving conformational exploration and accelerating convergence relative to simulations initiated only from endpoint structures. These results establish sequence-guided structural interpolation as a practical strategy for probing protein conformational landscapes. By generating diverse and physically reasonable conformational proposals without long-timescale MD or model retraining, Pi-Ensemble provides an extensible framework for studying protein flexibility, guiding adaptive sampling, and accelerating mechanistic investigations of protein function.
Hassan Nadeem, D. Kleiman, Yuming Zhou et al.· bioRxiv· 0 citations