Skip to content
Open access

A Machine Learning-Based Genome Mining Approach Reveals Unprecedented Biarylitide Diversity

Jul 2026 · JACS Au · Vol 6, pp. 3952 - 3965 · 0 citations · 48 references
Medicine

TL;DR

This study repurposed a machine learning algorithm to comprehensively chart the biosynthetic space of the biarylitides, including variation of precursor motifs, P450, and additional modifying enzymes, which yielded 277 biarylitide biosynthetic gene clusters (BGCs).

Abstract

Biarylitides are a group of bacterial ribosomally synthesized and post-translationally modified peptides (RiPPs) that contain a biaryl bridge formed by dedicated cytochrome P450 enzymes that can introduce different cross-links. The biarylitides are produced via a five-amino-acid precursor peptide, encoded by a minimal 18 bp gene that evades automatic detection. Previous genome mining approaches for biarylitides do not capture their full biosynthetic space. We therefore repurposed a machine learning algorithm to comprehensively chart the biosynthetic space of the biarylitides, including variation of precursor motifs, P450, and additional modifying enzymes, which yielded 277 biarylitide biosynthetic gene clusters (BGCs). We experimentally investigated biaryl formation with previously uninvestigated core peptide motifs, including YWH, YVH, and YWY, and elucidated the nature of these cross-links. This study significantly expands the biarylitide precursor and BGC diversity and provides directions for the systematic exploration of other RiPP families.

Read PDF

Similar papers

Open access Jul 2026

Discovery of a phenazine–thiol conjugase from sparse data using genome-informed machine learning

Machine learning has enabled powerful biological discoveries using models trained on large datasets. However, for many important biological questions, such as identifying enzymes that transform understudied substrates, sparsity of training data is often a major bottleneck. Here, using phenazine natural products as a case study, we show that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences. We name this approach ML-CITO (Machine Learning for genomic Context-Informed Transferable discOvery). Applying this framework led to the discovery of PTC (Phenazine-Thiol Conjugase), the first enzyme known to catalyze phenazine thioconjugation, a phenazine modification reaction long observed but previously presumed to occur only through non-enzymatic chemistry. In silico simulation and experimental measurements demonstrate that PTC binds to both phenazine and glutathione as substrates. Recombinant expression and biochemical characterization reveal that PTC promotes glutathione-dependent modification of phenazines, yielding distinct reaction outcomes that depend on substrate identity. Although thiol-conjugated phenazine products exhibit reduced toxicity to bacterial cells, deletion of the gene encoding PTC does not confer a strong fitness disadvantage, illustrating how direct learning of sequences can uncover relevant enzymes that might evade phenotype-based genetic screens. Together, these results demonstrate that coupling comparative genomics with protein machine learning can convert “small data” typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery.

Xiaoyu Shan, I. Trindade, N. Glasser et al. · 0 citations
Open access Aug 2026

Combining Machine Learning and Directed Evolution for Optimization of a Monooxygenase

A systematic comparison of zero-shot ML models is provided and an iterative framework for integrating machine learning with directed evolution to accelerate enzyme engineering is established to accelerate enzyme engineering.

Daniel Gutierrez, Isa Madrigal Harrison, Aaron L. Feller et al. · 0 citations
Open access Dec 2025

DeepAden: an explainable machine learning method for predicting the substrate specificity of nonribosomal peptide synthetases

DeepAden achieves competitive performance compared with state-of-the-art tools on a benchmark dataset, and enabled the identification of two Streptomyces NRPS gene clusters through accurate A-domain substrates specificity predictions.

Jiaquan Huang, Liangjun Ge, Yaxin Wu et al. · 0 citations
Jul 2026

Machine-Learning-Enabled Rapid Evolution of Photoenzymes for the Asymmetric Synthesis of gem-Difluorophosphonates.

A small-sample, accelerated evolution strategy that integrates focused rational iterative site-specific mutagenesis (FRISM) with the EVOLVEpro model is reported, providing a robust, "lightweight" machine learning framework for the rapid development of new-to-nature photoenzymatic transformations.

Hongkui Wang, Jiafan Xu, Jiahai Zhou et al. · 0 citations
Open access Jul 2026

Machine Learning-Assisted Evolution of Broadly Functional Enzyme Libraries

Results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope, and suggest that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope.

Ravi G. Lal, Jason Yang, Ziyan Zhang et al. · 0 citations
Open access Aug 2026

Computational mass spectrometry and genome mining guided discovery of metallophores produced by Microbulbifer

Iron is an essential component of cellular biology. Thus, iron's low bioavailability is a key evolutionary pressure guiding microbial dynamics in the marine environment. Among marine bacteria, Microbulbifer is a chemically underexplored and functionally versatile bacterial genus, which is commonly associated with sponges, algae, corals and sediments. Previously, genome analyses have revealed that Microbulbifer spp. can degrade polymers and synthesize natural products. Despite their recognized potential to produce secondary metabolites, siderophores are yet to be identified in Microbulbifer, and their iron acquisition strategies remain largely unknown. Here, we developed a comprehensive mass spectrometry-based query language (MassQL) code to determine siderophore production by Microbulbifer spp. in mono- and mixed cultures. Using this workflow, we discovered a new metallophore, which we named bulbichelin, as well as a suite of previously unreported petrobactins containing an unprecedented longer chain length acylation on the central spermidine moiety. We applied genome mining methods to describe the biosynthesis of these compounds. Using metal infusion mass spectrometry, we show that bulbichelins bind a variety of metals. Notably, neither of these compounds were produced in a co-culture of Microbulbifer with coral-derived pathogen Vibrio coralliilyticus Cn52-H1. Understanding how siderophores shape interspecies interactions between Microbulbifer spp. and other marine organisms will aid in unraveling the chemical and catalytic versatility of this genus and adaptation in nutrient deplete marine environment.

Mónica Monge-Loría, C. Brady, Hongwei Wu et al. · 0 citations