Mycosporine-like amino acids (MAAs) are functional secondary metabolites renowned for their exceptional UV protection, antioxidant properties, and environmental resilience. In the MAA biosynthetic pathway, MysD is the pivotal enzyme mediating the chemical transition of the cyclohexenone core into a cyclohexenimine-type scaffold. This MysD-catalyzed secondary amino acid modification not only dictates the chemical diversity of MAAs but also facilitates a crucial bathochromic shift, moving the UV absorption maximum from the UVB range into the high-penetration UVA region. Despite its significance, MysD remains the rate-limiting step in the biosynthesis of iminomycosporine-like amino acids. To date, only eight MysD enzymes have been heterologously validated, with even fewer subjected to biochemical characterization, which severely restricts the use of conventional supervised machine learning for enzyme discovery. In this study, we developed an integrated, data-driven screening framework combining Sequence Similarity Networks (SSN), deep representation learning (UniRep), and Positive-Unlabeled Bagging (PU Bagging) to explore the MysD functional landscape. This pipeline effectively compressed the search space from approximately 951 unannotated homologues to a prioritized 42 candidates. Experimental validation led to the discovery of AdMysD from Aphanothece hegewaldii, which exhibited a 3-fold increase in catalytic efficiency for porphyra-334 production relative to the previously established benchmark, NlMysD. Notably, while demonstrating a primary preference for l-Thr, AdMysD displayed significant substrate promiscuity by accepting l-Ser, l-Ala, and l-Cys to produce iminomycosporine derivatives. Our findings provide a biocatalytic tool for the efficient production of MAAs and demonstrate the potential of a PU-learning-based prioritization strategy for identifying rare enzyme families with sparse functional annotations.
A small-sample, accelerated evolution strategy that integrates focused rational iterative site-specific mutagenesis (FRISM) with the EVOLVEpro model is reported, providing a robust, "lightweight" machine learning framework for the rapid development of new-to-nature photoenzymatic transformations.
This study repurposed a machine learning algorithm to comprehensively chart the biosynthetic space of the biarylitides, including variation of precursor motifs, P450, and additional modifying enzymes, which yielded 277 biarylitide biosynthetic gene clusters (BGCs).
Leo Padva, Jemma Gullick, Friederike Biermann et al.· JACS Au· 0 citations
A transfer learning-based predictor for acidophilic and alkalophilic proteins, trained on a curated non-redundant dataset, and an integrated pipeline for mining and engineering alkalophilic and thermophilic enzymes, combining sequence-based prediction, generative modeling, and multi-parameter virtual screening are developed.
Ruohan Zhang, Yiyang Zhang, Zhonghao Deng et al.· Bioresource Technology· 0 citations
The rapid emergence of metallo-b-lactamase-mediated antibiotic resistance has created an urgent need for new inhibitor discovery strategies. In this work, a machine-learning-guided workflow was developed to generate and prioritize potential inhibitors targeting NDM-1. A SMILES-based variational autoencoder was first pretrained on a broad molecular dataset to learn general chemical syntax and latent molecular representations. The model was then fine-tuned on an 8-hydroxyquinoline-enriched dataset to bias molecular generation toward zinc-binding chemical space relevant to metallo-β-lactamase inhibition. Generated compounds were processed through structural filtering and docking-based evaluation to create training data for downstream predictive modeling. Molecular fingerprints and physicochemical descriptors were then used to train XGBoost models for docking score prediction and classification of potential binders. Classification proved especially useful for prescreening because it avoided overinterpreting small differences in noisy docking scores while still enriching for compounds likely to perform well in docking. The resulting workflow demonstrates how generative modeling and supervised machine learning can be combined to reduce chemical search space, prioritize candidate inhibitors, and guide computational drug discovery. Although experimental validation remains necessary, this approach provides a scalable framework for identifying promising zinc-binding compounds for further molecular simulation and inhibitor development that can be expanded in future studies.
Anthony M. Baudino, Kari L. Stone· AI Chemistry· 0 citations
This work demonstrates that integrating cell-free enzyme engineering with machine learning enables opportunities for high-throughput experimental measurements to benchmark and improve protein language models, accelerate design loops, and expand functional exploration within protein families where experimental information is limited.
J. Lazar, Evan Komp, I. Martínez et al.· bioRxiv· 1 citation
This study engineered an experimental callus system for inducible production of CPT, which enabled multi‐omics and deep learning analyses to identify candidate genes in CPT biosynthesis and provides a valuable foundation for the complete elucidation of the CPT biosynthetic pathway.
Shenqiu Wang, Xing Wu, Maria Moreno et al.· The Plant Genome· 0 citations