Skip to content
Open access

Discovery of non-canonical proteins through modification-aware proteogenomics

Aug 2026 · bioRxiv · 0 citations · 56 references
Biology

TL;DR

The potential for open modification searching to correct potential mistakes in non-canonical proteins detection by preventing modified canonical peptides or variants from being incorrectly identified as non-canonical peptides is shown.

Abstract

Short The SwissProt database contains a stable 20,418 human protein-coding genes and 42,541 human protein sequences. Ribo-Seq suggests about 7,000 additional, non-canonical Open Reading Frames (ORFs) are present in humans, though only a few of them are confirmed by Mass Spectrometry (MS). Detecting these proteins requires extensive database searches, increasing computational load and inflating False Discovery Rates (FDR). Using the ionbot search engine with the OpenProt database allows for reliable detection of non-canonical proteins while controlling FDR. Ionbot surpasses the Trans-Proteomics Pipeline (TPP) in reproducibility, identifying more peptides and proteins supported by multiple spectra. In addition, open modification searches yield better PSMs compared to closed searches. This work highlights the importance of employing cutting-edge search engines in non-canonical protein research, as well as the value of open modification search in correcting errors in non-canonical protein detection. Long Background The SwissProt database reports a quite stable 20,418 human protein-coding genes and 42,541 human protein sequences, figures that have remained stable. New techniques like Ribo-Seq indicate that approximately 7,000 additional, non-canonical Open Reading Frames (ORFs) are translated in humans, few of which have been confirmed by Mass Spectrometry (MS). Detecting these non-canonical proteins requires comprehensive database searches, which increase computational load and False Discovery Rate (FDR). Here, we use the open search engine ionbot in combination with the OpenProt proteogenomics database to reproducibly detect non-canonical proteins while maintaining a well-controlled FDR. Results Compared to the current gold standard, the Trans-Proteomics Pipeline (TPP), ionbot shows higher reproducibility, with a higher number of peptides and proteins supported by multiple spectra, and across multiple samples. We observe that PSMs from the open modification search against OpenProt have higher fragment ion intensity correlation compared to PSMs obtained from the closed search, or by only searching canonical proteins. Conclusions In this work, we show the potential for open modification searching to correct potential mistakes in non-canonical proteins detection by preventing modified canonical peptides or variants from being incorrectly identified as non-canonical peptides. We also highlight the importance of assessing the FDR of non-canonical identifications separately from canonical ones, as global FDR calculations are biased by the scarcity of non-canonical identifications in each dataset.

Read PDF

Similar papers

Review Open access Aug 2026

XMAn Update - A Database of Homo sapiens Mutated Peptides

Mass spectrometry (MS) is the leading technology for identifying proteins in complex biological samples. It relies on the use of tandem MS alongside a reference database of canonical protein sequences to computationally identify peptides and their parent proteins. The canonical sequences represent the most widely expressed and functionally validated forms of proteins. Consequently, disease-induced or disease-supportive variants, such as those associated with cancer, will evade detection if they are absent from the database. To address this challenge, this study introduces a revised release of the Unkown Mutation Analysis (XMAn) database by incorporating coding missense and nonsense mutations from the latest versions (v103) of the COSMIC Genome Screen Mutants (GSM) and Cancer Gene Census (CGC) datasets in two distinct FASTA-formatted peptide databases comprising 3,848,499 and 312,658 variants, respectively. The mutated peptides were matched to reviewed, non-redundant UniProt Homo sapiens protein entries (18,362 and 746), and characterized in terms of nucleotide- and amino acid mutation frequencies, peptide length distributions, and associations between specific single-nucleotide (SNV) and single amino acid (SAAVs) variants. Applied to the analysis of MDA-MB-231 breast cancer cell-membrane protein fractions, the database enabled the identification of 300+ high-quality variant peptides - several localized to functional protein-binding and catalytic domains - and 23 aberrant protein products mapped to the CGC dataset. The database is hosted and available for download on Zenodo (XMAn/gsm doi: 10.5281/zenodo.21781023; XMAn/cgc doi: 10.5281/zenodo.21781514) or can be accessed through https://sites.google.com/vt.edu/xman-db/home.

Joshua R. S. Haueis, I. Lazar · 0 citations
Open access Aug 2026

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Guillermo Carrillo-Martin, Johanna Krueger, T. Marquès-Bonet et al. · 0 citations
Open access Aug 2026

Mapping the Human Ghost Proteome: Classification and Experimental Detection Biases in the Identification of Alternative Microproteins

The discovery of alternative proteins (AltProts), translated from non-canonical ORFs, has expanded the human proteome and revealed a hidden layer known as the “ghost proteome”. Despite increasing evidence, AltProts detection remains challenging due to their small size, physicochemical heterogeneity, and lack of annotation. Here, we developed an integrated bioinformatic and proteomic workflow to benchmark the detection of reference proteins (RefProts), isoforms, and alternative microproteins (MicroAltProts) in colorectal cancer cells using four extraction protocols—HCl, RIPA buffer, RIPA with chloroform, and RIPA followed by 30 kDa filtration—combined with high-resolution data-independent acquisition mass spectrometry. We identified and quantified using the Orbitrap Astral mass spectrometer a total of 66,438 peptides corresponding to 12,584 different protein groups across methods, with RIPA-based extraction approaches providing the most comprehensive coverage. To reduce redundancy in the OpenProt database and focus on MicroAltProts, we curated the dataset by removing known isoforms and long proteins, yielding a non-redundant set of 183,937 MicroAltProts. K-means clustering based on eight ProtParam-derived features grouped MicroAltProts into four physicochemical clusters. Among them, 43 MicroAltProts (<200 amino acids) were experimentally validated by mass spectrometry and classified into tiers following recent recommended international guidelines. Cluster assignment of detected MicroAltProts revealed that HCl extraction favored disordered, alkaline proteins, while RIPA-based protocols enabled the identification of membrane-associated and amphipathic α-helical MicroAltProts. Structural prediction indicated the presence of diverse folding determinants, including transmembrane helices, disordered regions, and nucleic acid-binding-like motifs. Altogether, this study provides a roadmap framework for the unbiased simultaneous detection of RefProts, isoforms, and AltProts, and supports a broader functional role for MicroAltProts.

A. Montero‐Calle, A. Peláez-García, A. J. Martín-Galiano et al. · 0 citations
Open access Aug 2026

GenomeProt: User friendly proteogenomics for canonical and non-canonical proteoform characterisation

Quantifying the diversity of RNAs and proteins produced by cells is fundamental to the biological and clinical sciences. However, many RNAs and proteins remain uncharacterised, especially proteins translated from alternate RNA isoforms; untranslated regions of mRNAs and non-coding RNAs, as well as the effects of DNA variation on protein sequences. Proteogenomics aims to characterise the complete proteome by integrating genomics and/or transcriptomics with proteomics, but current tools have limitations in useability, analysis features and visualisation of resulting data. To address these gaps, we developed GenomeProt, a user-friendly GUI-based tool for integrative proteogenomic analysis. We demonstrate its utility by integrating long-read RNA sequencing with mass-spectrometry-based proteomics to pinpoint proteoform expression generated by alternative splicing; discover novel, unannotated proteins in human brain samples; and quantify variant-containing peptides associated with treatment resistance in a melanoma xenograft model. GenomeProt brings the discovery power of proteogenomics to biologists, illuminating the hidden proteome.

Hitesh Kore, Josie Gleeson, C. Wan et al. · 1 citation
Open access Aug 2026

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.

Diya Mathai, S. Schulze · 0 citations
Open access Sep 2026

Background proteome correction promotes confident identification of dynamic protein-protein interactions between different biological contexts

Affinity purification-mass spectrometry (AP-MS) enables the characterization of protein-protein interactions (PPIs), and the ease and sensitivity of such experiments has progressively increased. Beyond steady-state interactions of target proteins, a strong interest has emerged in monitoring how PPIs change upon significant biological perturbations, such as in disease contexts or small molecule modulation of the target protein. These perturbations likely not only induce PPI changes but can also lead to altered expression of proteins not of direct interest. Changes in protein abundance may alter which proteins adsorb to the affinity purification matrix, and due to the sensitivity of modern mass spectrometers, these differential “background binders” can masquerade as differential interactors. Contemporary approaches often do not account for differences in the background proteome, potentially inflating the number of false positives and negatives reported. Here, we provide technical considerations for the reliable annotation of dynamic PPIs, using the O-GlcNAc transferase (OGT) as a case study. We describe the installation of affinity epitope tags on endogenous OGT in mouse embryonic stem cells (mESCs), which we then apply for OGT interactor identification via AP-MS. We show that accurate representation of the bead background, which depends on the affinity matrix in use, is critical for elimination of false positive and false negative PPIs. This became even more pertinent as OGT PPI dynamics were measured under OGT catalytic inhibition via OSMI-4, which is known to perturb gene expression. The proteomes of OSMI-4-treated and control-treated mESCs differed, leading to distinct bead backgrounds in which the differential background proteins appeared as interaction gains or losses. These false positives were resolved by incorporating straightforward experimental controls through a practical statistical framework, allowing for a direct and confident comparison between treatment conditions. Incorporating these considerations into workflows investigating PPI dynamics will improve data fidelity and reproducibility.

M. Brunelli, Lisa Morishita-Cartwright, Natalie M. Clark et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.