The Genomic Annotation Infrastructure (GAIn) is presented, a platform that generates transparent, reproducible annotations via declarative pipelines that define annotation tasks as ordered lists of components that produce annotation attributes using genomic resources from Genomic Resource Repositories (GRRs).
Abstract
Interpretation of genomic variants, positions, and regions depends on reliable annotation—adding evidence such as predicted effect, conservation, population frequency, and gene-level context—yet the underlying resources are numerous, versioned, and assembly-specific. We present the Genomic Annotation Infrastructure (GAIn), a platform that generates transparent, reproducible annotations via declarative pipelines that define annotation tasks as ordered lists of components, called annotators, that produce annotation attributes using genomic resources from Genomic Resource Repositories (GRRs). We provide two public GRRs: a main repository containing more than 250 heterogeneous genomic resources, and a separate GRR-ENCODE repository containing resources derived from thousands of ENCODE (Encyclopedia of DNA Elements) project experiments. Users can use the annotation pipelines we made available, author custom annotation pipelines, and execute annotation tasks with these pipelines via GAIn’s web and command-line interfaces. The web interface can be used without any setup, but it relies on shared computational infrastructure and imposes limits on the size of annotation tasks. The command-line interface requires setup but supports arbitrarily large annotation tasks through simple-to-use parallelization and offers a broader set of features. For example, command-line GAIn can be extended by using custom GRRs or creating custom annotators via its plugin architecture. In addition, GAIn’s re-annotation feature, which updates annotations as they evolve, substantially simplifies maintaining annotations in a large genomics analysis project. GAIn’s resource management, explicit versioning, and pipeline abstraction provide an auditable, maintainable, and efficient foundation for modern genomic annotation across reference assemblies and use cases.
The rapid growth in publicly available genome assemblies has made selecting genomes suitable for downstream analyses increasingly challenging. Differences in assembly and annotation quality can influence gene completeness, duplication rates, contiguity, repeat representation, and other characteristics. Assessing genome...
Christopher D. R. Wyatt, Fernando Duarte Frutos, Stephen D. Turner et al.· bioRxiv· 0 citations
Abstract Summary Gene annotation of metagenome-assembled genomes is a critical step in determining the functional potential of microbial communities from environmental samples. However, annotation workflows using tools such as Prokka or Bakta produce per-bin output with 10 to 14 files per bin, making manual review infe...
Kepler Ridge, Byron J. Adams· Bioinformatics Advances· 0 citations
Background
Prokaryotic genome annotation is central to comparative genomics, functional interpretation, and hypothesis generation. Established tools such as Prokka and Bakta provide streamlined annotation workflows, but predefined database choices and hierarchical annotation strategies can limit flexibility, especially...
Richard Stöckl, Felix Grünberger, Dina Grohmann· F1000Research· 0 citations
GEOMeta provides a scalable resource and reproducible framework for metadata curation in the Gene Expression Omnibus, and benchmarked transcriptome representation models for predicting sex, age, tissue and disease from transcriptome embeddings.
Xiaodan Zhang, S. Paithankar, Jing Pu et al.· bioRxiv· 0 citations
Scop3P-Toolkit is an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers.
Adrián Díaz, Natalia Tichshenko, Boris Depoortere et al.· bioRxiv· 0 citations
High-throughput sequencing has generated protein datasets whose scale increasingly exceeds the practical limits of conventional functional annotation workflows. We present Sma3s v3, a scalable reimplementation of the Sma3s three-step annotation strategy, which combines transfer from highly similar homologs, orthology-b...
Alejandro Rubio, Jesús L. García-Junco Alcalá, Elisa Luque-Jiménez et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.