Skip to content
Open access

Pandoomain, a scalable pipeline for genomic and protein domain context analysis, reveals widespread PT-TG domain architectural diversity and novel polymorphic toxins

Jul 2026 · mSystems · Vol 11 · 0 citations · 58 references
Medicine

TL;DR

Pandoomain is an accessible tool that enables systematic, large-scale exploration of protein domains, and the analysis of the PT-TG domain provides a rich resource for future investigations into the mechanisms and evolution of bacterial antagonism.

Abstract

ABSTRACT The rapid expansion of bacterial genome databases presents significant opportunities for functional discovery, as a large fraction of genes and protein domains remain uncharacterized. Analyzing genomic context and domain architecture is a powerful approach for functional inference, but existing tools often lack the scalability and integrated workflow required for high-throughput analysis. To address this, we developed Pandoomain, a Snakemake pipeline that automates the acquisition of genomes from the National Center for Biotechnology Information, identifies proteins of interest using hidden Markov models (HMMs), and performs systematic domain annotation and gene neighborhood analysis. We demonstrate the utility of Pandoomain through a comprehensive analysis of the poorly characterized pre-toxin TG (PT-TG) domain across 347,289 bacterial genomes. Our analysis revealed 10,226 PT-TG-containing proteins organized into 312 unique domain architectures, highlighting their association with diverse interbacterial antagonistic systems, including the Type VI secretion, Type VII secretion, and contact-dependent inhibition systems. By leveraging genomic context, we identified a novel variant of the WXG trafficking domain, termed W10XG, and subsequently discovered 24 new families of associated toxin domains. We experimentally validated six of these toxins, confirming that all six are neutralized by their cognate immunity proteins. Pandoomain is an accessible tool that enables systematic, large-scale exploration of protein domains, and our analysis of the PT-TG domain provides a rich resource for future investigations into the mechanisms and evolution of bacterial antagonism. IMPORTANCE The rapid growth of bacterial genomic data presents a major hurdle for scientists seeking to understand the functions of newly discovered genes and proteins. To address this issue, we created Pandoomain, a powerful, accessible software tool that automates large-scale analysis of genetic information across hundreds of thousands of genomes. Using Pandoomain, we investigated a poorly understood family of proteins involved in bacterial competition, revealing novel protein domain architectural diversity. This led to the discovery of 24 new families of toxins predicted to be used by bacteria to attack their competitors, and we experimentally confirmed the toxic activity of six of them. Our work provides the scientific community with a robust tool to accelerate functional discovery and offers new insights into the evolution of bacterial conflicts, which may provide insights into the compositional dynamics of microbial communities and support methods to engineer their composition. The rapid growth of bacterial genomic data presents a major hurdle for scientists seeking to understand the functions of newly discovered genes and proteins. To address this issue, we created Pandoomain, a powerful, accessible software tool that automates large-scale analysis of genetic information across hundreds of thousands of genomes. Using Pandoomain, we investigated a poorly understood family of proteins involved in bacterial competition, revealing novel protein domain architectural diversity. This led to the discovery of 24 new families of toxins predicted to be used by bacteria to attack their competitors, and we experimentally confirmed the toxic activity of six of them. Our work provides the scientific community with a robust tool to accelerate functional discovery and offers new insights into the evolution of bacterial conflicts, which may provide insights into the compositional dynamics of microbial communities and support methods to engineer their composition.

Read PDF

Similar papers

Open access Jul 2026

PhytoFam: A Nextflow Pipeline for Genome-Wide Analysis of Plant Gene Families

Genome-wide identification of plant gene families is essential for functional and evolutionary studies but often requires the use of multiple independent tools for homolog detection, domain validation, orthology assignment, and phylogenetic analysis. This fragmented approach involves extensive manual scripting, complic...

Sanam Parajuli, Bibek Adhikari, Anne Y. Fennell et al. · 0 citations
Open access Aug 2026

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-s...

Diya Mathai, S. Schulze · 0 citations
Jul 2026

A Scalable Pan-Genomic Pipeline for Annotation-Free Discovery of Species-Specific Markers: Application to Staphylococcus aureus

An open-source, annotation-independent pan-genomic pipeline featuring an overlapping sliding-window algorithm and a three-tier subtractive screening funnel successfully circumvents conventional gene-centric limitations, providing a generalizable computational strategy for target discovery across other high-priority bac...

Yu-Yang Zhou, Jiayi Wang, Jun-Hua Xiao et al. · 0 citations
Open access Jul 2026

BGX: A Comprehensive Pipeline for Genomic Insight into Bioactivity Prediction, Genomic Surveillance, and Novel Biosynthetic Gene Cluster Assessment

The modular and reproducible architecture of the BGX pipeline provides an effective framework for large-scale genome mining, genomic surveillance, and accelerates the discovery and prioritisation of novel secondary metabolites.

Ardhendu Chakrabortty, Lovepreet Singh, Babanpreet Kaur et al. · 0 citations
Open access Aug 2026

Scop3P-Toolkit: executable structure-aware workflows linking PTMs, peptides, and mutations to protein function

Scop3P-Toolkit is an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers.

Adrián Díaz, Natalia Tichshenko, Boris Depoortere et al. · 0 citations
Open access Aug 2026

ppigFinder: an integrated desktop application for bacterial genome annotation and AlphaFold 3-based protein–protein interaction screening

ppigFinder combines ORF prediction, functional annotation, genomic-neighbourhood inspection, AlphaFold 3 job generation, remote job submission, and structural-confidence analysis within a single environment for genome-based PPI discovery from nucleotide sequence data.

G. U. Oka, Camilla Adan, Celso Vítor Alves Queiroz Calomeno et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.