Skip to content
Open access

An in-depth update on the benchmarks for 16S amplicon sequencing

Aug 2026 · bioRxiv · 0 citations · 35 references
Biology

TL;DR

This study conducted a comprehensive benchmark of the main bioinformatic tools and databases and demonstrated that even certain dominant genera remain difficult to detect, and that the quantification of all genera can be substantially over- or under-estimated, even when using optimal combinations of tools and reference databases.

Abstract

Amplicon-based techniques provide a rapid and cost-effective approach for profiling microbial communities. However, the observed microbial diversity is influenced by a wide range of factors, encompassing pre-analytical steps such as the choice of primers and target regions, as well as the bioinformatic pipeline, including the selection of tools, reference databases, and parameter settings. Several benchmarks are already available in the literature, but the updates to important tools and databases, namely LotuS3, the Ribosomal Database Project and GreenGenes2, prompted our investigation. In this study, we conducted a comprehensive benchmark of the main bioinformatic tools and databases. Using seven regions for three publicly available mock communities of increasing complexity, we tested 38 possible combinations of sequence resolution algorithms (DADA2 stand-alone, LotuS3 (DADA2/UPARSE)), taxonomic classifiers and search tools (Kraken2, DECIPHER, RDP, MMseqs2, Lambda, and Metaxa2), and databases (SILVA, GreenGenes2, RDP, RefSeq, and Metaxa2). The region V1-V3, coupled with DADA2+MMseqs2+SILVA, DADA2+Metaxa2, or LotuS3 (DADA2)+RDP yielded the highest-quality estimates of the true diversity according to the metrics. We also demonstrated that even certain dominant genera remain difficult to detect, and that the quantification of all genera can be substantially over- or under-estimated, even when using optimal combinations of tools and reference databases.

Read PDF

Similar papers

Review Open access Aug 2026

Bioinformatic tools for microbiome analysis: from raw sequences to biological insights

This review presents a practical, workflow-oriented guide to microbiome data analysis, from raw DNA sequence processing to statistical interpretation and biological insight, and highlights emerging technologies, including machine learning methods that are beginning to reshape the field.

Jenna Poelzer, D. Wishart · 0 citations
Open access Aug 2026

When bigger is not better: the impact of RefSeq growth on Kraken2 classification accuracy

This study analyzed the influence of RefSeq database size and composition on taxonomic identification performance using Kraken 2, a widely used taxonomic classification and profiling method and found that resource-efficient Kraken 2 Lite (capped) databases generally exhibit lower classification accuracy compared to ful...

Leonardo Lazzaro, Enrico Rossignolo, Matteo Comin · 0 citations
Sep 2026

A bioinformatics framework using public 16S rRNA gene amplicon data to assess the presence of target bacteria in bat and rodent samples.

A dual-strategy bioinformatics pipeline that leverages publicly available 16S rRNA gene amplicon sequencing data to reliably and inexpensively confirm target bacterial presence and distinguished target-positive from negative samples, with phylogenetic support for specificity is described.

Jian Zhou, T. Gu, Shi-Jun Li · 0 citations
Open access Sep 2026

Large-scale benchmarking of prokaryotic annotation tools across thousands of species

This study presents the first large-scale investigation of four prominent open-source annotation tools (Prokka, Bakta, EggNOG-mapper, and PGAP) across 156,033 diverse genomes, highlighting tool-specific strengths crucial for selecting optimal solutions based on genome quality, taxonomy, and origin.

Mateusz Jundzill, Martin Hölzer, S. Mangul et al. · 0 citations
Aug 2026

maniFasta and the AllOralsDB: simplifying construction of comprehensive reference databases for metaproteomics

ManiFasta is presented, a tool enabling users to generate standardized, reproducible, and robustly documented protein reference sets from diverse input sources and datatypes, and the value of maniFasta is highlighted in the context of salivary metaproteomics, addressing the need for a taxonomically comprehensive refere...

Christopher Handelmann, Ashley K. Miles, Yin-Yin Ye et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.