Skip to content
Review Open access

nf-core/magmap: Map metatranscriptomes to large collections of genomes

Jul 2026 · Bioinformatics · Vol 42 · 0 citations · 27 references
Medicine

TL;DR

The nf-core/magmap pipeline is presented, which provides a reproducible, easy-to-access, and well-documented workflow for selecting reference genomes, mapping to them, and quantifying features.

Abstract

Abstract Summary The lack of publicly available reference genomes has forced annotation of metatranscriptomes to either use direct alignment of sequence reads to reference databases or de novo assembly. As more and more natural environments are covered by metagenomic surveys, this is rapidly changing. This opens up the possibility of genome-resolved studies of prokaryotic metatranscriptomes by mapping to genomes from public repositories or metagenome-assembled genomes derived from the same environment. Here, we present the nf-core/magmap pipeline that provides a reproducible, easy-to-access, and well-documented workflow for selecting reference genomes, mapping to them, and quantifying features. Genomes can be drawn from public sources or originate from private collections. The pipeline is primarily aimed at prokaryotic communities but can, together with collections of reference mature gene sequences, also be applied to eukaryotes. Availability and implementation The nf-core/magmap pipeline is implemented in Nextflow and part of the nf-core collaboration. The pipeline is available at the nf-core website (https://nf-co.re/magmap) and GitHub (https://github.com/nf-core/magmap).

Read PDF

Similar papers

Open access Aug 2026

GenomeCompendium: A database for the integrated analysis of repeats, assembly quality and functional content of complete prokaryotic genomes

Microorganisms hold great promise for urgent global needs such as increasing sustainable agricultural production while reducing chemical fertilizer and pesticide use or providing novel classes of antimicrobials/therapeutics. Moving from analyzing microbiome composition to applying synthetic communities and studying their functions requires access to isolates and complete genome sequences. By spanning the frequent repeats, long-read sequencing can resolve complex prokaryotic genomes, yet error-prone short-read assemblies dominate. We here release the GenomeCompendium, a public database and interactive analysis tool for complete prokaryotic genomes (https://genome-compendium.com/). Using NCBI RefSeq (∼47,000) and GenBank (∼13,000) genomes, we integrated available metadata, GTDB taxonomy and computed features including repeat classification and frequency analysis, intragenomic 16S rRNA sequence identity, and biosynthetic gene cluster co-occurrences. Evaluating repeat content and assembly complexity metrics, we identify taxonomic ranks dominated by difficult-to-assemble genomes and show that complex, repeat-rich genomes are more common than previously estimated. By mining metadata, our quality control flags 6.3% of RefSeq assemblies as potentially erroneous or incomplete. As valuable reference for data mining and to track taxonomic coverage, the GenomeCompendium links ∼90 features across genomes, offers downloadable reports and -as unique features-pre-computed proteogenomics databases to improve genome annotations of RefSeq strains and the ability to analyze any uploaded prokaryotic genome.

Tiberiu Totu, Garance Jaques, B. Heiniger et al. · 0 citations
Open access Jul 2026

Annotation of glycoside hydrolases in unassembled metagenomes using CAZyOGH

Abstract Motivation Functional characterization of microbiomes often relies on the sequencing of metagenomic DNA extracted from environmental samples, with current approaches using metagenome-assembled genomes (MAGs). Although glycoside hydrolases (GHs) are central to carbon cycling, accurate annotation of GHs in metagenomic datasets remains challenging due to the multidomain architecture of carbohydrate-active enzymes and the prevalence of unassembled short reads due to limitations in the MAG-generation process. Results Here, we present CAZyOGH (CAZymes Open-source GH annotation), a curated reference database for the domain-specific identification of 135 protein domains spanning 99 GH families with well-defined catalytic domain signatures. CAZyOGH focuses on individual GH domains, enabling robust annotation of both assembled and unassembled metagenomic data. We validated CAZyOGH by reanalyzing genomes listed in CAZy db, where predicted GH profiles closely matched reported values. Next, we used CAZyOGH to analyze 12 human gut metagenomes and 12 newly sequenced soil microbiomes to reveal environment-specific GH repertoires. By accurately detecting catalytic domains independent of the genomic context, CAZyOGH improves sensitivity and specificity in short-read metagenomic annotation. This framework provides a scalable and reproducible approach to investigate carbohydrate-active enzymes across ecosystems, advancing our capacity to characterize microbial functional potential in global carbon cycling. Availability and implementation CAZyOGH data is available on figshare (https://figshare.com/projects/CAZyO_GH/267770).

N. Griffin, Alison E Hughes, D. S. Erdody et al. · 0 citations
Open access Aug 2026

metaIVP: an integrative metavirome focused metagenomic processing pipeline

Metagenomic studies increasingly rely on complex, multi-tool pipelines to recover and characterize viral and non-viral genomes from mixed microbial communities. While these pipelines enable high-resolution genome recovery, limited functionality in downstream post-processing workflows and insufficient logging structures often hinder reproducibility, error tracing, and selective re-analysis. These challenges are particularly critical in metaviral analyses, where viral and non-viral genomes must be processed using distinct methodologies. To address these limitations, we introduce metaIVP, a modular, integrative, and flexible framework designed to systematically manage genome content purification, re-binning, quality assessment, and downstream analyses of viral and non-viral metagenomic contexts. The metaIVP framework is organized into hierarchical modules, each governed by dedicated log files that explicitly control execution state and re-runnability. Contig-level and bin-level analytical and purification steps are implemented as essential modules to isolate genome contents, followed by separate viral and non-viral post-processing workflows. Viral workflows incorporate contamination detection, genome quality evaluation, host prediction, and virus-specific binning. Non-viral analyses include genome binning, alignment and mapping statistics, genome quality assessment, and replication rate estimation. Checkpoints are explicitly defined such that deletion of selected module- or sub-module-level logs enables targeted re-execution of specific analytical steps without rerunning the full pipeline. All analyses are integrated to depict a comprehensive system in the metagenomic samples, with focus on the metaviromic information. The usage of metaIVP was demonstrated using both a well-controlled human gut virome dataset and a geographically structured environmental metavirome dataset, showing its broad applicability across host-associated and environmental systems. The pipeline effectively separates viral and non-viral genomic content, improves viral bin purity, and preserves sample-specific functional, taxonomic, and host-association features after virome enrichment. Compared with recent state-of-the-art approaches, metaIVP achieves comparable performance, particularly when optional re-binning with vRhyme is applied, while maintaining a higher fraction of high-confidence viral bins. The metaIVP addresses a key gap in metavirome analysis by jointly characterizing viral and non-viral genomic components and supporting integrative downstream analyses within a single framework. Its user-friendly, modular, and controllable design allows flexible execution and provides a foundation for incorporating additional downstream analytical tools as metavirome methodologies continue to evolve.

Kalyan Sahu, Qiuming Yao · 0 citations
Open access Aug 2026

nf-core/genomeqc: a best-practice pipeline for comparing genome and assembly quality

The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome quality therefore requires integrating multiple complementary quality metrics that are often generated by independent tools. Here, we present nf-core/genomeqc, a workflow for assessing and comparing genome assemblies. The pipeline accepts RefSeq/GenBank accessions for automatic genome and annotation retrieval, or local genome (FASTA) and annotation (GFF3/GTF) files. It integrates complementary analyses of assembly contiguity, gene completeness, annotation quality, repeat content and other quality metrics using tools such as BUSCO, QUAST, Merqury, and AGAT, before combining the results on a phylogenetic tree for visualisation and comparison across species. GenomeQC is implemented in Nextflow within the nf-core framework, providing an accessible, reproducible, scalable and community-driven workflow for genome quality assessment.

C. D. Wyatt, Fernando Duarte Frutos, Stephen D. Turner et al. · 0 citations
Review Open access Aug 2026

FAIRyMAGs - a series of FAIR Galaxy workflows for the generation of metagenome assembled genomes

Advances in whole-genome sequencing (WGS) technologies have enabled large-scale recovery of metagenome-assembled genomes (MAGs), providing unprecedented insights into microbial diversity across diverse environments. However, the reconstruction of MAGs remains computationally demanding and methodologically complex, requiring the integration of multiple tools for quality control, assembly, binning, refinement, and annotation. Existing workflows often rely on scripting-based implementations, constrain user-driven modification and stepwise execution, and require advanced expertise in high-performance computing (HPC) system administration, thereby limiting accessibility, reproducibility, and adaptability. Here, we present FAIRyMAGs, a Findable, Accessible, Interoperable, and Reusable (FAIR)-compliant, modular pipeline implemented within the Galaxy platform for the generation and analysis of MAGs. FAIRyMAGs consists of six interconnected workflows covering all major steps of MAG reconstruction, including read preprocessing, host and contaminant removal, assembly, binning, dereplication, and downstream taxonomic and functional annotation. The workflows are accompanied by extensive training material, including tutorials, a learning pathway, FAQs, test datasets and video walk-throughs by domain experts, supporting community adaptation. By leveraging Galaxy’s graphical interface and federated infrastructure, FAIRyMAGs enables users to execute complex analyses on public or private compute resources without requiring local installation or workflow programming expertise. The modular design further supports flexible adaptation, iterative optimization, and seamless integration of new tools contributed by the community. To demonstrate applicability, FAIRyMAGs was applied to four real-world microbiome datasets spanning various host-associated and environmental systems. These analyses revealed substantial variability in MAG recovery, community complexity, and clustering structure, underscoring the importance of flexible workflows adaptable to dataset-specific characteristics. Overall, FAIRyMAGs provides an accessible, extensible, and reproducible framework for genome-resolved metagenomics, reducing technical barriers and enabling methodological innovation through community-driven development within the adaptable Galaxy ecosystem.

P. Zierep, Mina Hojat Ansari, Patrick Bühler et al. · 0 citations