This study conducted a comprehensive benchmark of the main bioinformatic tools and databases and demonstrated that even certain dominant genera remain difficult to detect, and that the quantification of all genera can be substantially over- or under-estimated, even when using optimal combinations of tools and reference databases.
Abstract
Amplicon-based techniques provide a rapid and cost-effective approach for profiling microbial communities. However, the observed microbial diversity is influenced by a wide range of factors, encompassing pre-analytical steps such as the choice of primers and target regions, as well as the bioinformatic pipeline, including the selection of tools, reference databases, and parameter settings. Several benchmarks are already available in the literature, but the updates to important tools and databases, namely LotuS3, the Ribosomal Database Project and GreenGenes2, prompted our investigation. In this study, we conducted a comprehensive benchmark of the main bioinformatic tools and databases. Using seven regions for three publicly available mock communities of increasing complexity, we tested 38 possible combinations of sequence resolution algorithms (DADA2 stand-alone, LotuS3 (DADA2/UPARSE)), taxonomic classifiers and search tools (Kraken2, DECIPHER, RDP, MMseqs2, Lambda, and Metaxa2), and databases (SILVA, GreenGenes2, RDP, RefSeq, and Metaxa2). The region V1-V3, coupled with DADA2+MMseqs2+SILVA, DADA2+Metaxa2, or LotuS3 (DADA2)+RDP yielded the highest-quality estimates of the true diversity according to the metrics. We also demonstrated that even certain dominant genera remain difficult to detect, and that the quantification of all genera can be substantially over- or under-estimated, even when using optimal combinations of tools and reference databases.
This review presents a practical, workflow-oriented guide to microbiome data analysis, from raw DNA sequence processing to statistical interpretation and biological insight, and highlights emerging technologies, including machine learning methods that are beginning to reshape the field.
Jenna Poelzer, D. Wishart· Frontiers in Microbiology· 0 citations
This study analyzed the influence of RefSeq database size and composition on taxonomic identification performance using Kraken 2, a widely used taxonomic classification and profiling method and found that resource-efficient Kraken 2 Lite (capped) databases generally exhibit lower classification accuracy compared to ful...
Leonardo Lazzaro, Enrico Rossignolo, Matteo Comin· Frontiers in High Performanc...· 0 citations
This review revisits algorithms, tools, and workflows for sequence and phylogenetic analysis in the NGS-based omics era, with a focus on comparative performance and scenario-driven decision-making.
Abhishek Kumar, T. C. Dakal, Kayenat Parveen et al.· Biochemical Genetics· 0 citations
A dual-strategy bioinformatics pipeline that leverages publicly available 16S rRNA gene amplicon sequencing data to reliably and inexpensively confirm target bacterial presence and distinguished target-positive from negative samples, with phylogenetic support for specificity is described.
Jian Zhou, T. Gu, Shi-Jun Li· Journal of Microbiological M...· 0 citations
This study presents the first large-scale investigation of four prominent open-source annotation tools (Prokka, Bakta, EggNOG-mapper, and PGAP) across 156,033 diverse genomes, highlighting tool-specific strengths crucial for selecting optimal solutions based on genome quality, taxonomy, and origin.
Mateusz Jundzill, Martin Hölzer, S. Mangul et al.· Genome Biology· 0 citations
ManiFasta is presented, a tool enabling users to generate standardized, reproducible, and robustly documented protein reference sets from diverse input sources and datatypes, and the value of maniFasta is highlighted in the context of salivary metaproteomics, addressing the need for a taxonomically comprehensive refere...
Christopher Handelmann, Ashley K. Miles, Yin-Yin Ye et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.