Sep 2026· Journal of Microbiological Methods· pp.
107684
· 0 citations
Medicine
TL;DR
A dual-strategy bioinformatics pipeline that leverages publicly available 16S rRNA gene amplicon sequencing data to reliably and inexpensively confirm target bacterial presence and distinguished target-positive from negative samples, with phylogenetic support for specificity is described.
Abstract
Validating the ecological distribution of a newly isolated bacterial species in natural hosts remains challenging due to the lack of specific detection assays and the cost of large-scale screening. Here, we describe a dual-strategy bioinformatics pipeline that leverages publicly available 16S rRNA gene amplicon sequencing data to reliably and inexpensively confirm target bacterial presence. The method first extracts hypervariable regions from the target bacterium's full-length 16S rRNA gene and evaluates their specificity by calculating an A-value-defined as the highest sequence similarity to any non-target strain in reference databases. Regions with an A-value below the 98.7% species threshold are selected. These are then aligned against Amplicon Sequence Variants (ASVs) from public datasets to compute a B-value (highest similarity to ASVs within a sample). A novel classification logic (B > A) is applied to designate samples as positive or negative, reducing false positives. The pipeline incorporates multi-level controls, including process/biological negatives and positives. Testing with novel species (Clostridium sp. nov.) and a formally described species (Streptococcus lishijunsis), along with common commensal species demonstrated that region-specific performance varies, highlighting the need for pre-validation. The framework successfully distinguished target-positive from negative samples, with phylogenetic support for specificity. This approach provides a rigorous, cost-effective, and accessible workflow that links in vitro isolation to in vivo ecological validation using existing public data.
Rapid and accurate microbial identification is critical for interpreting biological data in basic research and when making applied decisions on how to effectively treat patients and control human, animal, and plant diseases. Advancements in high-throughput sequencing have the potential to expedite fungal species identi...
Hayden Johnson, B. Vinatzer, R. Mazloom et al.· bioRxiv· 0 citations
Determining which members of a microbial community are metabolically active remains a central challenge in microbial ecology. Although the 16S rRNA gene is the dominant marker for bacterial community profiling, it cannot reliably distinguish active cells from dormant or dead populations. As a result, complementary phyl...
Fabien Cholet, William T. Sloan, Cindy J. Smith· bioRxiv· 0 citations
Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford...
Ida Romano, E. Pasolli, J. Walser et al.· BMC Genomics· 0 citations
Background:Antimicrobial resistance (AMR) poses an important challenge to public health on a global scale, with traditional methods of susceptibility testing not sufficiently fast to allow empirical treatment or surveillance. Whole-genome sequencing (WGS) coupled with machine learning (ML) offers a promising, genome-sc...
GA Al-Oudah, Nada Khazal K. Hindi, I. Abdul-Husin et al.· Journal of Biomedicine and B...· 0 citations
Species-resolved PCR detection of gut microorganisms is challenged by within-species genomic diversity and sequence conservation among related taxa. We developed PrimerBac (https://primerbac.cn), a searchable database that integrates gut-focused reference collection, RefSeq pan-genome analysis, core-orthogroup marker s...
Ying-Jian Hou, Shuo Feng, Mao-Yuan Gu et al.· Journal of Molecular Biology· 0 citations