Skip to content
Open access

GenomeCompendium: A database for the integrated analysis of repeats, assembly quality and functional content of complete prokaryotic genomes

Aug 2026 · bioRxiv · 0 citations · 81 references
Biology

Abstract

Microorganisms hold great promise for urgent global needs such as increasing sustainable agricultural production while reducing chemical fertilizer and pesticide use or providing novel classes of antimicrobials/therapeutics. Moving from analyzing microbiome composition to applying synthetic communities and studying their functions requires access to isolates and complete genome sequences. By spanning the frequent repeats, long-read sequencing can resolve complex prokaryotic genomes, yet error-prone short-read assemblies dominate. We here release the GenomeCompendium, a public database and interactive analysis tool for complete prokaryotic genomes (https://genome-compendium.com/). Using NCBI RefSeq (∼47,000) and GenBank (∼13,000) genomes, we integrated available metadata, GTDB taxonomy and computed features including repeat classification and frequency analysis, intragenomic 16S rRNA sequence identity, and biosynthetic gene cluster co-occurrences. Evaluating repeat content and assembly complexity metrics, we identify taxonomic ranks dominated by difficult-to-assemble genomes and show that complex, repeat-rich genomes are more common than previously estimated. By mining metadata, our quality control flags 6.3% of RefSeq assemblies as potentially erroneous or incomplete. As valuable reference for data mining and to track taxonomic coverage, the GenomeCompendium links ∼90 features across genomes, offers downloadable reports and -as unique features-pre-computed proteogenomics databases to improve genome annotations of RefSeq strains and the ability to analyze any uploaded prokaryotic genome.

Read PDF