Testing on a large cohort of soil metagenomes, it is found that BBERT identifies bacterial sequence syntax without relying on reference databases, enabling accurate assignment of taxonomic domain, coding potential, and reading frame directly from reads as short as 100 bp.
Abstract
Shotgun metagenomes from complex environments such as soil uncover vast biodiversity. Yet most short reads produced by shotgun sequencing cannot be taxonomically or functionally annotated, as they lack a sufficiently comprehensive reference, obscuring the true structure and function of microbial communities. We introduce BBERT, a nucleotide large language model optimized for short reads. Testing on a large cohort of soil metagenomes, we found that BBERT identifies bacterial sequence syntax without relying on reference databases, enabling accurate assignment of taxonomic domain, coding potential, and reading frame directly from reads as short as 100 bp. BBERT is small and fast enough to analyze metagenomes using a modest GPU and can be used to convert short metagenomic reads directly to bacterial amino acid sequences for downstream applications. BBERT also improves de-novo metagenomic assembly, reducing mismatches and gaps while accelerating runtime. Using metagenomes from wild legume nodules, we demonstrate that BBERT filtering improves bin quality while significantly accelerating de-novo assembly. By providing fast, reference-free classification of short reads, BBERT unlocks large metagenomic archives for more accurate ecological and evolutionary analyses. Graphical Abstract
ABSTRACT Functional profiling of meta-omics is essential for understanding microbial communities, yet support for custom genome-resolved reference databases is limited. We introduce Leviathan for integrated taxonomic and functional profiling at both genome and pangenome resolution. Leviathan combines Sylph for ultrafas...
Josh L. Espinoza, Allan J. Phillips, C. Dupont· mSystems· 0 citations
This study presents the first large-scale investigation of four prominent open-source annotation tools (Prokka, Bakta, EggNOG-mapper, and PGAP) across 156,033 diverse genomes, highlighting tool-specific strengths crucial for selecting optimal solutions based on genome quality, taxonomy, and origin.
Mateusz Jundzill, Martin Hölzer, S. Mangul et al.· Genome Biology· 0 citations
The results establish Sma3s v3 as a scalable and interpretable tool for functional annotation and re-annotation of proteomes, pangenomes, and metagenomic protein catalogues.
Alejandro Rubio, Jesús L. García-Junco Alcalá, Elisa Luque-Jiménez et al.· bioRxiv· 0 citations
Generating high-quality genome assemblies for small animals with large genomes is complex due to their small body size, DNA contamination, and repetitive elements. Ticks exemplify these complexities, while also being a global health threat to humans, domestic animals, and wildlife. Advances in long-read sequencing plat...
Katie C. Dillon, H. Sprong, Isobel Ronai et al.· bioRxiv· 0 citations
Abstract Summary Gene annotation of metagenome-assembled genomes is a critical step in determining the functional potential of microbial communities from environmental samples. However, annotation workflows using tools such as Prokka or Bakta produce per-bin output with 10 to 14 files per bin, making manual review infe...
Kepler Ridge, Byron J. Adams· Bioinformatics Advances· 0 citations
Background/Objectives: Oxford Nanopore sequencing produces long reads quickly, but most functional profiling tools were developed for short reads or rely on assembly pipelines that are computationally costly and sensitive to long-read error rates. We present NanoPrism, a taxonomy-guided pipeline for rapid functional pr...
Jiwoong Kim, Shuheng Gan, Harish Jawahar et al.· DNA· 0 citations