antiSMASH is widely used for biosynthetic gene cluster (BGC) detection and annotation, but its standard workflow is poorly suited to large metagenomic assemblies, where massive contig counts create severe runtime bottlenecks and complicate downstream result exploration. We present metaSMASH, a re-engineered fork of antiSMASH for metagenome-scale BGC analysis. metaSMASH preserves the original antiSMASH detection and annotation logic while introducing streaming, memory-bounded execution, record-level parallelisation, optional output filtering, and an interactive dashboard for large result sets. Across 25 benchmark metagenome datasets, metaSMASH reproduced identical BGC detection results while dramatically reducing computational cost. Relative to the default antiSMASH configuration, metaSMASH was a geometric-mean 38× faster. It also outperformed an ad hoc chunked antiSMASH workflow: in the default configuration it achieved a geometric-mean 2.9 × speed-up and 1.7 × lower peak memory, and with extended-analysis modules enabled it was 2.7 × faster and used 3.1 × less memory while completing all datasets, whereas the ad hoc workflow ran out of memory on the two largest assemblies. By substantially reducing the computational burden of large-scale metagenome analysis without sacrificing result equivalence, metaSMASH makes routine mining of assembled metagenomes more practical and provides a scalable foundation for natural product discovery from complex microbial communities. Graphical Abstract
The results establish Sma3s v3 as a scalable and interpretable tool for functional annotation and re-annotation of proteomes, pangenomes, and metagenomic protein catalogues.
Alejandro Rubio, Jesús L. García-Junco Alcalá, Elisa Luque-Jiménez et al.· bioRxiv· 0 citations
This work introduces orthoSynAssign, a user-friendly, high-performance rewrite of the orthogroup refinement tool OrthoRefine, combining an intuitive Python interface with a core computing engine written in Rust, which provides a reliable and scalable framework for high-throughput phylogenomic workflows.
Cheng-Hung Tsai, Carolina G. Piña Páez, J. Stajich· bioRxiv· 0 citations
The metaIVP is introduced, a modular, integrative, and flexible framework designed to systematically manage genome content purification, re-binning, quality assessment, and downstream analyses of viral and non-viral metagenomic contexts that addresses a key gap in metavirome analysis.
FAIRyMAGs provides an accessible, extensible, and reproducible framework for genome-resolved metagenomics, reducing technical barriers and enabling methodological innovation through community-driven development within the adaptable Galaxy ecosystem.
P. Zierep, Mina Hojat Ansari, P. Bühler et al.· bioRxiv· 0 citations
PyiTOL validates inputs, generates 31 iTOL template schemas, performs LCA-based monophyly classification with nested-monophyly detection, sampling-completeness states and polyphyletic subgroup decomposition, plus API upload and session replay.
This work presents REAPER (Repeatome Extended Analysis Pipeline—Execution and Reporting), a project-centric workflow layer that couples a modular Snakemake pipeline with a Python project manager to enforce a stable on-disk layout and configuration-driven execution for single-sample and comparative repeatome analyses.
D. Ulyanov, A. I. Yurkina, Viktoria Voronezhskaya et al.· Frontiers in Plant Science· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.