The ChlORIS database is presented, presenting a large-scale plastid phylogeny of 859 taxa spanning all major algal lineages and characterising the protein HMMs by cross-referencing them to Pfam domains and calibrated score cutoffs for reliable detection.
Abstract
Chloroplast or plastid genomes are essential resources for studying the evolution and diversity of algae and land plants. Although thousands of plastid genomes have been sequenced, their full potential has not been realised; derived resources such as orthogroup databases and reference datasets for metagenomic profiling remain underdeveloped. We present the ChlORIS database to address these problems across all algal phyla. From 2,254 publicly available algal plastid genomes, after dereplication we clustered 2,531 orthogroups from the annotated proteins and selected 496 orthogroups with consistent gene naming, enabling cross-genome comparisons of homologous plastid proteins. We further selected 224 core orthogroups, each containing more than 10 protein sequences, for which we produced score-calibrated hidden Markov models (HMMs), multiple sequence alignments and predicted protein structures. The value of these resources for phylogenomics is demonstrated through a large-scale plastid phylogeny of 859 taxa spanning all major algal lineages. We characterised the protein HMMs by cross-referencing them to Pfam domains and calibrated score cutoffs for reliable detection. The metagenomic database, HMM library, nucleotide and amino acid alignments, predicted structures and protein metadata, cross-linked to UniProt and InterPro (Pfam), are openly available on the ChlORIS website at https://chloris.codeberg.page/.
Introduction Primula is the largest genus in Primulaceae and a classic model for studying heterostyly and plant evolution. Methods In this study, we sequenced, assembled, and annotated the complete chloroplast genomes of seven Primula species from China. Two additional plastomes of Primula dejuniana and Primula meishan...
Yan-Ru Zhang, Shi-Hao Jiang, Qin-Yi Wang et al.· Frontiers in Plant Science· 0 citations
Pangenome analysis reveals recurrent gene-content variation beyond a single reference genome, but its application to eukaryotes is constrained by inconsistent gene annotation. ANNEVO predicts gene models from genome FASTA assemblies without RNA-seq data. We developed EukPan, an automated post-annotation pipeline that s...
K. Seki, Masatoshi Goto, Taiki Futagami et al.· bioRxiv· 0 citations
The GenomeCompendium is released, a public database and interactive analysis tool for complete prokaryotic genomes and it is shown that complex, repeat-rich genomes are more common than previously estimated.
Tiberiu Totu, Garance Jaques, B. Heiniger et al.· bioRxiv· 0 citations
The SAR11 Genome Atlas is presented, an interactive ortholog group (OG)-centered web resource that integrates 542 SAR11 genomes, including all 132 cultured strain genomes, with functional annotations, synteny, phylogenetic distribution, metatranscriptomic expression, and predicted protein structure information.
Tardigrades are microscopic invertebrates renowned for their capacity to survive extreme environmental conditions, yet stress tolerance strategies vary substantially across the phylum. Here we present the first genome assembly and annotation for the superfamily Isohypsibioidea — the most basal group within the order Pa...
Daria S. Makarova, Denis Tumanov, Elena Nassonova et al.· Functional & Integrative Gen...· 0 citations
New methodologies were examined, including genome scanning, advanced assembly tools such as GetOrganelle, and multispecies merger phylogenetic reconstruction, highlighting the necessity of multi-genome integration, the application of pan-plastome methodologies, and the expanding possibilities of chloroplast synthetic b...
Shaima Abdul Rahman, M. Karaismailoğlu· Bartın University Internatio...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.