Aug 2026· Science Advances· Vol 12· 0 citations· 72 references
Medicine
TL;DR
FEDKEA, an enzyme annotation tool leveraging protein language models, and a user-friendly, FEDKEA-based metagenomic pipeline, MEnzMap, which encompasses the entire analysis workflow—from raw data quality control to function prediction and downstream analyses are designed.
Abstract
Metagenomic data have notable biological potential, but their functional interpretation is frequently impeded by incomplete protein function annotations. Accurate enzyme annotation is essential for elucidating the metabolic capabilities of microbial communities within metagenomic datasets. To address this challenge, we developed FEDKEA, an enzyme annotation tool leveraging protein language models, and provided a web platform for its use. In addition, we designed a user-friendly, FEDKEA-based metagenomic pipeline, MEnzMap, which encompasses the entire analysis workflow—from raw data quality control to function prediction and downstream analyses. Applying MEnzMap to human gut metagenomic data from the iHMP2 project, we generated a comprehensive enzyme profile landscape for both healthy individuals and patients with inflammatory bowel diseases. These tools provide an efficient method for the functional annotation of microbial dark matter and facilitate the identification of disease-associated enzymes.
The rapid growth of microbiome research has been accompanied by an expanding but fragmented ecosystem of bioinformatic tools. Researchers now face a daunting array of software packages, pipelines, and web platforms spanning every stage of analysis, from quality control and taxonomic profiling to functional annotation and statistical interpretation. While this diversity offers flexibility, it also creates challenges in selecting appropriate tools and integrating them into coherent, reproducible workflows, particularly for researchers without formal computational training. This review presents a practical, workflow-oriented guide to microbiome data analysis, from raw DNA sequence processing to statistical interpretation and biological insight. We evaluate tools based on ease of use, methodological rigor, computational requirements, and community support, with particular attention to the trade-offs between command-line interface and web-based approaches. We cover both amplicon and shotgun metagenomic strategies for taxonomic and functional profiling, discuss reference database selection, and outline key statistical methods, including differential abundance testing and network inference. We also compare integrated platforms and web-based resources that lower barriers for non-computational researchers and discuss best practices for reproducibility and workflow design. Throughout, we highlight emerging technologies, including machine learning methods that are beginning to reshape the field. Overall, this review serves as a practical guide to navigating the microbiome bioinformatics landscape, helping bridge the gap between methodological complexity and the biological questions that drive microbiome research.
Jenna Poelzer, D. Wishart· Frontiers in Microbiology· 0 citations
GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.
Xiaodan Zhang, S. Paithankar, Jing Pu et al.· bioRxiv· 0 citations
Protein function annotation is crucial for understanding biological processes and mechanisms. Traditionally, annotations rely on sequence homology, providing valuable insights but often leaving gaps even in well-characterised organisms. With AlphaFold enabling rapid generation of protein structural models, we can now infer function from three-dimensional shape. Here, we present WASP, a pipeline leveraging structural homology to enhance protein annotation prediction at scale, providing a more comprehensive understanding of protein functions across various organisms. WASP relies on network topology for better accuracy and more robust statistical power. We show that WASP achieves superior F1 scores compared to state-of-the-art sequence-based tools when recovering hidden annotations. On 20 industrially relevant organisms, WASP retrieves annotations for 20-30% of previously uncharacterised proteins. We further demonstrate utility in genome-scale metabolic model curation, identifying native candidates for 75-100% of orphan reactions. WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches. WASP predicts protein functions from AlphaFold structures using network-based structural homology, retrieving annotations for 20-30% of uncharacterised proteins and filling metabolic model gaps by mapping 75-100% of orphan reactions.
Abstract Motivation Functional characterization of microbiomes often relies on the sequencing of metagenomic DNA extracted from environmental samples, with current approaches using metagenome-assembled genomes (MAGs). Although glycoside hydrolases (GHs) are central to carbon cycling, accurate annotation of GHs in metagenomic datasets remains challenging due to the multidomain architecture of carbohydrate-active enzymes and the prevalence of unassembled short reads due to limitations in the MAG-generation process. Results Here, we present CAZyOGH (CAZymes Open-source GH annotation), a curated reference database for the domain-specific identification of 135 protein domains spanning 99 GH families with well-defined catalytic domain signatures. CAZyOGH focuses on individual GH domains, enabling robust annotation of both assembled and unassembled metagenomic data. We validated CAZyOGH by reanalyzing genomes listed in CAZy db, where predicted GH profiles closely matched reported values. Next, we used CAZyOGH to analyze 12 human gut metagenomes and 12 newly sequenced soil microbiomes to reveal environment-specific GH repertoires. By accurately detecting catalytic domains independent of the genomic context, CAZyOGH improves sensitivity and specificity in short-read metagenomic annotation. This framework provides a scalable and reproducible approach to investigate carbohydrate-active enzymes across ecosystems, advancing our capacity to characterize microbial functional potential in global carbon cycling. Availability and implementation CAZyOGH data is available on figshare (https://figshare.com/projects/CAZyO_GH/267770).
N. Griffin, Alison E Hughes, D. S. Erdody et al.· Bioinformatics Advances· 0 citations
The functional complexity inherent in microbiomes complicates analytical approaches aimed at defining phenotype. As proteins are the functional effectors of microbiome phenotypes, improving the performance of mass spectrometry-based metaproteomics is critical to achieving the functional characterization of these systems. Data-independent acquisition (DIA) improves protein coverage and reduces data missingness when compared to data-dependent acquisition (DDA) in metaproteomics. However, the application of DIA to complex microbial systems remains constrained by analytical throughput and computational scalability. Here, we optimized LC–MS/MS acquisition parameters for both DDA and DIA using a model microbiome, demonstrating how DIA enables increased sample throughput without compromising quantitative performance. In addition, we demonstrated a computationally efficient, library-free DIA workflow that overcomes reliance on empirical spectral libraries. Our analytical and computational innovations establish a scalable and cost-effective pipeline for metaproteomics of complex microbial communities.
Samantha Obermiller, Mary S. Lipton, P. Piehowski et al.· bioRxiv· 0 citations