Dec 2025· bioRxiv· Vol 54· 0 citations· 78 references
BiologyMedicine
TL;DR
DeepAden achieves competitive performance compared with state-of-the-art tools on a benchmark dataset, and enabled the identification of two Streptomyces NRPS gene clusters through accurate A-domain substrates specificity predictions.
Abstract
Microbial non-ribosomal peptides (NRPs) exhibit remarkable structural diversity and serve as valuable sources of lead compounds for clinical drug development. The biosynthesis of NRPs relies on non-ribosomal peptide synthetases (NRPSs), in which adenylation (A) domains play a pivotal role in defining the core structure by selectively recognizing and activating amino acid substrates. Accurately predicting the substrate specificities of A-domains is thus essential for understanding the core structural and biosynthetic logic of NRPs. Here, we present DeepAden, a two-stage deep learning framework. In the first stage, a graph attention network (GAT)-based model localizes 27-residue binding pockets within 6 Å of bound substrates and convert these into pocket representations. In the second stage, pocket representations are then encoded alongside substrate information using pretrained language models, and aligned using contrastive learning. In addition, we introduce a SHapley Additive exPlanations (SHAP)-guided data augmentation strategy to mitigate class imbalance and improve robustness, particularly for nonproteinogenic substrates. DeepAden achieves competitive performance compared with state-of-the-art tools on a benchmark dataset, and enabled the identification of two Streptomyces NRPS gene clusters through accurate A-domain substrates specificity predictions. DeepAden offers a powerful tool for precise pocket localization and robust substrate prediction, accelerating the discovery and characterization of novel NRP natural products for future work. The DeepAden web server is available at https://deepnp.site/.
This study repurposed a machine learning algorithm to comprehensively chart the biosynthetic space of the biarylitides, including variation of precursor motifs, P450, and additional modifying enzymes, which yielded 277 biarylitide biosynthetic gene clusters (BGCs).
Leo Padva, Jemma Gullick, Friederike Biermann et al.· JACS Au· 0 citations
An unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets is presented and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.
Nazmul Shuzan, Jialun Wei, Jie Zheng· Journal of Chemical Informat...· 0 citations
Accurate computational prediction of enzyme function, standardized by Enzyme Commission (EC) numbers, is essential for large-scale genome annotation and generative enzyme design. However, it remains unclear whether state-of-the-art predictors learn the intrinsic structural determinants of catalytic activity or merely rely on global sequence similarity to annotated homologues. To address this gap, we introduce EnzymARC, a novel benchmark dataset of putative non-functional decoy sequences generated via structure-guided, systematic disruption of active sites (targeting catalytic residues and surrounding 5 Å, 10 Å, and 15 Å radii) from experimentally annotated enzymes. We evaluated three distinct prediction paradigms against this dataset: homology-based annotation (DIAMOND), contrastive learning with protein language models (CLEAN), and a deep learning model incorporating non-enzyme discrimination (DeepEC). Our findings reveal that current models are highly vulnerable to phylogenetic shortcuts. Both DIAMOND and CLEAN exhibited false positive rates exceeding 90% for low-perturbation decoys, confidently assigning the original EC numbers despite the destruction of the catalytic machinery. While DeepEC demonstrated improved sensitivity at higher perturbation levels—highlighting the benefit of negative training examples—all models struggled to identify targeted active-site disruptions. We demonstrate that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and we propose that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.
João Sartori, Ana Carolina Ramos Guimarães, Lucas de Almeida Machado· bioRxiv· 0 citations
Introduction Microbial small proteins, encoded by small open reading frames (smORFs), play essential roles in antimicrobial activity, metabolic regulation, and signaling pathways. However, their short length and rapid evolutionary rate present significant challenges for computational modeling. Methods We introduce TinyProteinTransformer (TPT), a lightweight CNN-Transformer hybrid encoder pretrained on the Global Microbial smORF Catalog (GMSC; >280 million smORF families). TPT integrates multi-scale convolutional filters to capture local sequence motifs with Transformer layers for broader contextual modeling, and is trained jointly with masked language modeling and contrastive learning to capture both residue-level and sequence-level representations. We evaluated TPT on six downstream peptide/protein classification benchmarks spanning antimicrobial peptides (AMP), toxic peptides (TOX), bacteriocins (BCN), anti-CRISPR proteins (Acr), quorum-sensing peptides (QSP), and cell-penetrating peptides (CPP) using frozen-encoder linear probing. Results On the AMP and TOX tasks, TPT (103 M parameters) matched ESM2-150 M in predictive performance (AUC 0.930 vs. 0.928 for AMP; 0.930 vs. 0.925 for TOX) while achieving 4.3-fold faster inference. Across the evaluated benchmarks, TPT showed competitive performance compared with larger protein language models. Ablation experiments identified the contrastive objective as an important contributor to representation quality: its removal reduced AUC across all six downstream tasks and produced performance patterns consistent with MLM-only pretraining. Conclusion These results suggest that, within the evaluated benchmarks, effective smORF representation may benefit from pretraining objectives and inductive biases tailored to short, rapidly evolving sequences, rather than from model scale alone. TPT therefore provides a compact and computationally efficient encoder with competitive performance for microbial peptide analysis.
BACKGROUND
Internal ribosome entry sites (IRES) are cap-independent translation initiation elements present in specific viral and cellular mRNAs. They facilitate direct ribosome recruitment for protein synthesis, bypassing the need for the canonical 5' cap structure. Due to their pivotal roles in viral pathogenesis and cellular translational regulation, precise identification of IRES is crucial for advancing mechanistic studies and exploring potential therapeutic interventions. Nonetheless, manual identification is both labor-intensive and costly, while current computational methods exhibit limitations in accuracy and robustness.
RESULTS
DB-IRES is a novel deep learning model designed for this purpose, incorporating densely connected 1D convolutional neural network blocks, bi-directional gated recurrent units, and a self-attention mechanism within an ensemble learning framework. Trained with a five-fold cross-validation strategy, the final integrated model exhibits superior and more robust predictive performance compared to existing methods on an independent test set. The model consistently achieves enhanced discriminatory capabilities across multiple evaluation metrics, thereby validating the effectiveness of its hybrid architecture and ensemble design.
CONCLUSIONS
DB-IRES serves as a reliable and precise computational tool for predicting IRES elements. Its improved performance enables more in-depth functional investigations of IRES biology and supports wider applications in RNA research and the development of associated therapeutics.
Membrane molecular recognition features (MemMoRFs) are lipid-binding intrinsically disordered regions (IDRs) that undergo disorder-to-order transitions to mediate critical membrane dynamics. Consequently, their dysregulation is closely linked to severe human pathologies, including neurodegenerative diseases and viral infections. Despite their biological significance, annotations for MemMoRFs are scarce, limiting the accuracy of computational predictors. We introduce PreMemMoRF, a deep learning framework that leverages transfer learning to alleviate data scarcity. The model is pre-trained on linear interacting peptides (LIPs) with similar conformational transitions and fine-tuned on MemMoRF datasets, capturing generalizable binding-related sequence features. PreMemMoRF outperforms existing predictors across multiple metrics and demonstrates robust performance on transmembrane and membrane-associated proteins. It also performs consistently in short linear motif prediction, highlighting cross-task generalizability. Proteome-wide analysis in yeast shows that predicted scores exhibit systematic differences across distinct transmembrane topological regions and are consistent with established physicochemical constraints of membrane proteins. Collectively, these results validate PreMemMoRF as a robust and reliable computational framework for the large-scale identification of MemMoRFs.
Chenxi Xia, Jiayi Hao, Hao Liu et al.· IEEE journal of biomedical a...· 0 citations