A systematic comparison of zero-shot ML models is provided and an iterative framework for integrating machine learning with directed evolution to accelerate enzyme engineering is established to accelerate enzyme engineering.
Abstract
L-3,4-dihydroxyphenylalanine (L-Dopa) is an important pharmaceutical for the treatment of Parkinson’s disease and a precursor to numerous catechol-containing compounds. The flavin-dependent monooxygenase HpaBC is a promising biocatalyst for microbial L-Dopa production but exhibits limited native activity toward L-tyrosine. Although structure-based machine learning (ML) models have become increasingly popular for protein engineering, relatively few studies have systematically compared their performance or evaluated their integration into iterative engineering workflows. Here, we benchmarked multiple ML models for their ability to predict activity enhancing mutations in HpaBC. Experimentally validated single mutants were used to seed combinatorial design with EVOLVEpro, generating progressively improved higher-order variants. We next evaluated how expanding the EVOLVEpro training set with directed evolution derived variants influenced combinatorial predictions and finally explored an expanded sequence space by allowing combinations of both machine learning derived and directed evolution derived mutations. This workflow produced HpaBC variants with substantially improved activity. Although incorporating directed evolution data substantially altered EVOLVEpro’s predicted mutational trajectories, both training strategies converged on variants with comparable activities, demonstrating that distinct regions of sequence space can yield similarly optimized enzymes. Together, these results provide a systematic comparison of zero-shot ML models and establish an iterative framework for integrating machine learning with directed evolution to accelerate enzyme engineering.
Results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope, and suggest that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope.
Ravi G. Lal, Jason Yang, Ziyan Zhang et al.· bioRxiv· 0 citations
A machine-learning-assisted enzyme-engineering (MLEE) workflow that adds substrate-specific functional information to htFuncLib through an initial screening and sequencing round that may bypass the need for transition-state models and reduce the effort required for obtaining high-activity variants.
Li Wan, Mahdi Bagherpoor Helabad, Lena Fraedrich et al.· bioRxiv· 0 citations
This study repurposed a machine learning algorithm to comprehensively chart the biosynthetic space of the biarylitides, including variation of precursor motifs, P450, and additional modifying enzymes, which yielded 277 biarylitide biosynthetic gene clusters (BGCs).
Leo Padva, Jemma Gullick, Friederike Biermann et al.· JACS Au· 0 citations
A small-sample, accelerated evolution strategy that integrates focused rational iterative site-specific mutagenesis (FRISM) with the EVOLVEpro model is reported, providing a robust, "lightweight" machine learning framework for the rapid development of new-to-nature photoenzymatic transformations.
The rapid emergence of metallo-b-lactamase-mediated antibiotic resistance has created an urgent need for new inhibitor discovery strategies. In this work, a machine-learning-guided workflow was developed to generate and prioritize potential inhibitors targeting NDM-1. A SMILES-based variational autoencoder was first pretrained on a broad molecular dataset to learn general chemical syntax and latent molecular representations. The model was then fine-tuned on an 8-hydroxyquinoline-enriched dataset to bias molecular generation toward zinc-binding chemical space relevant to metallo-β-lactamase inhibition. Generated compounds were processed through structural filtering and docking-based evaluation to create training data for downstream predictive modeling. Molecular fingerprints and physicochemical descriptors were then used to train XGBoost models for docking score prediction and classification of potential binders. Classification proved especially useful for prescreening because it avoided overinterpreting small differences in noisy docking scores while still enriching for compounds likely to perform well in docking. The resulting workflow demonstrates how generative modeling and supervised machine learning can be combined to reduce chemical search space, prioritize candidate inhibitors, and guide computational drug discovery. Although experimental validation remains necessary, this approach provides a scalable framework for identifying promising zinc-binding compounds for further molecular simulation and inhibitor development that can be expanded in future studies.
Anthony M. Baudino, Kari L. Stone· AI Chemistry· 0 citations