Skip to content
Open access

When bigger is not better: the impact of RefSeq growth on Kraken2 classification accuracy

Aug 2026 · Frontiers in High Performance Computing · Vol 4 · 0 citations · 24 references

TL;DR

This study analyzed the influence of RefSeq database size and composition on taxonomic identification performance using Kraken 2, a widely used taxonomic classification and profiling method and found that resource-efficient Kraken 2 Lite (capped) databases generally exhibit lower classification accuracy compared to full reference databases that require high-performance computing infrastructure.

Abstract

Current technologies allow for the sequencing of microbial communities directly from the environment without prior culturing. One of the major problems when analyzing a microbial sample is to taxonomically annotate its reads to identify the species it contains. In this study, we delve into an evaluation of how variations in the reference database impact classification performance. This aspect assumes paramount significance, particularly when considering the continuous evolution and expansion of the NCBI Reference Sequence database (RefSeq) over time. The aim of this study is to analyze the influence of RefSeq database size and composition on taxonomic identification performance using Kraken 2, a widely used taxonomic classification and profiling method. We found that resource-efficient Kraken 2 Lite (capped) databases, which can be run on standard laptops, generally exhibit lower classification accuracy compared to full reference databases that require high-performance computing infrastructure. The size of the RefSeq database and its growth affect the performance of the k-mer-based algorithms in addition to the computing resources need. We reported that changes in RefSeq reference database over time influenced the accuracy of metagenomic taxonomic classification, and training the classifier with more data, for example, new species, worsen the results especially on the most complex microbial communities.

Read PDF

Similar papers

Open access Aug 2026

An in-depth update on the benchmarks for 16S amplicon sequencing

This study conducted a comprehensive benchmark of the main bioinformatic tools and databases and demonstrated that even certain dominant genera remain difficult to detect, and that the quantification of all genera can be substantially over- or under-estimated, even when using optimal combinations of tools and reference...

Louis-Maël Guéguen, Alban Mathieu, Olivier Périn et al. · 0 citations
Open access Sep 2026

Whole genome similarity provides a rapid, robust framework for classification of fungal taxa from the genus rank to intraspecies variants

Rapid and accurate microbial identification is critical for interpreting biological data in basic research and when making applied decisions on how to effectively treat patients and control human, animal, and plant diseases. Advancements in high-throughput sequencing have the potential to expedite fungal species identi...

Hayden Johnson, B. Vinatzer, R. Mazloom et al. · 0 citations
Open access Sep 2026

Improving Metagenomics Classification with Kmask: Entropy-Based Masking of Low-Complexity Regions

Accurate taxonomic classification in metagenomics is often compromised by low-complexity sequences, which lead to chance matches that in turn cause sequences to be misclassified. Here we present Kmask, an entropy-based masking tool implemented for use either standalone or as part of Kraken [1,2] database construction,...

Yu-Chen Ge, Edward Z. Li, Harun Mustafa et al. · 0 citations
Open access Sep 2026

FuncAnnoClust Web Application for Analyzing Prokaryotic Genome Annotations Using Multivariate Statistics and Machine Learning

Comparative functional analysis of prokaryotic genomes is a key area of bioinformatics, but the practical implementation of such research often faces technical barriers. For example, automatic genome annotation using the widely used Rapid Annotations using Subsystems Technology platform is complicated by the fact that...

D. Gutnik, I. Petrushin, T. I. Belykh et al. · 0 citations
Open access Aug 2026

Assessment of the impact of manual curation in BioCyc

Introduction BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literat...

R. Caspi, B. Wilson-Mortier, L. Moore et al. · 0 citations
Sep 2026

A bioinformatics framework using public 16S rRNA gene amplicon data to assess the presence of target bacteria in bat and rodent samples.

A dual-strategy bioinformatics pipeline that leverages publicly available 16S rRNA gene amplicon sequencing data to reliably and inexpensively confirm target bacterial presence and distinguished target-positive from negative samples, with phylogenetic support for specificity is described.

Jian Zhou, T. Gu, Shi-Jun Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.