Abstract Motivation Functional characterization of microbiomes often relies on the sequencing of metagenomic DNA extracted from environmental samples, with current approaches using metagenome-assembled genomes (MAGs). Although glycoside hydrolases (GHs) are central to carbon cycling, accurate annotation of GHs in metagenomic datasets remains challenging due to the multidomain architecture of carbohydrate-active enzymes and the prevalence of unassembled short reads due to limitations in the MAG-generation process. Results Here, we present CAZyOGH (CAZymes Open-source GH annotation), a curated reference database for the domain-specific identification of 135 protein domains spanning 99 GH families with well-defined catalytic domain signatures. CAZyOGH focuses on individual GH domains, enabling robust annotation of both assembled and unassembled metagenomic data. We validated CAZyOGH by reanalyzing genomes listed in CAZy db, where predicted GH profiles closely matched reported values. Next, we used CAZyOGH to analyze 12 human gut metagenomes and 12 newly sequenced soil microbiomes to reveal environment-specific GH repertoires. By accurately detecting catalytic domains independent of the genomic context, CAZyOGH improves sensitivity and specificity in short-read metagenomic annotation. This framework provides a scalable and reproducible approach to investigate carbohydrate-active enzymes across ecosystems, advancing our capacity to characterize microbial functional potential in global carbon cycling. Availability and implementation CAZyOGH data is available on figshare (https://figshare.com/projects/CAZyO_GH/267770).
Background: Sebacina vermifera is a fungus that belongs to the Basidiomycota phylum (order Sebacinales). It has potential as a biofertilizer because it forms mutualistic relationships with many plant species, including orchids and other flowering plants. However, it is still not fully understood how this fungus plays a role in the soil/rhizosphere ecosystems where it occurs.
Methods: To profile S. vermifera-associated microbial community structures in relation to the different agroecological zone types, a multi-platform metagenomic sequencing approach was employed using sequencing technology platforms such as PacBio long read, Illumina short Read and Oxford Nanopore. Data processing for these metagenomic sequence assemblies included multiple steps including quality control using Trimmomatic and fastp, metagenome assembly with MEGAHIT and SPAdes, taxonomic profiling with Kraken2 and MetaPhlAn4, and functional annotation through EggNOG-mapper, KEGG Orthology, and CAZy databases. Network and comparative genomic analyses were also performed to characterise potential microbial interactions, as well as unique gene content.
Results: The results of metagenomic analyses showed that genes associated with phosphate solubilization (e.g., phytases, acid phosphatases), nitrogen fixation (e.g., nifH, nifD), production of siderophores, and the biosynthesis of indole-3-acetic acid were present. The association of S. vermifera with rhizosphere microbial networks increased the occurrence of interactions between nitrogen-fixing bacteria, arbuscular mycorrhizal fungi, and plant growth-promoting rhizobacteria. Unique effector proteins and secreted hydrolases were identified that were distinct from those of related fungal species. The field trials demonstrated a 34-42% increase in plant biomass, a 28% increase in phosphorus uptake, and a 19% decrease in applied chemical fertilizer.
Conclusion: With its rich repertoire of functional genes and beneficial interactions with other microorganisms, Sebacina vermifera represents a potential new source of biofertilizers for use in agriculture. The use of this fungus will result in greater crop yields, less dependence on chemical fertilizers, and healthier soils, thereby supporting the development of sustainable and climate-resilient agricultural systems.
Prasun Craven· Journal of Agricultural Digi...· 0 citations
GeneHunt2 is presented, a scalable framework for multidomain annotation of CAZymes that integrates curated HMM profiles from dbCAN and Pfam into a unified, deduplicated database, enabling systematic identification of both CAZy and non-CAZy domains.
The nf-core/magmap pipeline is presented, which provides a reproducible, easy-to-access, and well-documented workflow for selecting reference genomes, mapping to them, and quantifying features.
Danilo Di Leo, E. Nilsson, George Westmeijer et al.· Bioinformatics· 0 citations
FEDKEA, an enzyme annotation tool leveraging protein language models, and a user-friendly, FEDKEA-based metagenomic pipeline, MEnzMap, which encompasses the entire analysis workflow—from raw data quality control to function prediction and downstream analyses are designed.
Lei Zheng, Bowen Li, Siqi Xu et al.· Science Advances· 0 citations
This study reconstructed the first comprehensive pangenome of B. bifidum using 1,351 high-quality genomes, including metagenome-assembled genomes to identify species-specific genetic and functional features and identified significant gain-of-function events.
Emanuele Selleri, G. Longhi, C. Tarracchini et al.· Microbiome Research Reports· 0 citations
Eight complete Salinivibrio genomes from Pearse Lakes are generated using Oxford Nanopore long-read sequencing and seven putative depolymerases that form a single accessory cluster in 15% of strains are identified, showing that annotation-dependent approaches can overlook genomic diversity and divergent enzyme families in non-model organisms.
Crystal E. Young, Harrison O'Sullivan, Hussain Alattas et al.· Microbial Genomics· 0 citations