Scop3P-Toolkit is an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers.
Abstract
Post-translational modifications (PTMs) and genetic variants regulate protein function, signalling, and disease, but their interpretation requires integration of sequence annotations with structural, interaction, and biophysical context. Although resources such as Scop3P, UniProt, the Protein Data Bank, and AlphaFold provide extensive annotations and structural information, integrating these data into reproducible structure-aware analyses still requires custom scripting and manual coordination between multiple independent tools. To address this challenge, we developed Scop3P-Toolkit, an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context. The toolkit integrates protein annotation retrieval with structural mapping, residue interaction network analysis, comparative structural analysis, and residue-level biophysical profiling within a unified framework. Experimentally supported phosphosites, phosphopeptides, and phosphoproteomics evidence are provided for human proteins through Scop3P, with optional integration of curated UniProt PTM annotations. UniProt-derived PTMs, sequence features, and genetic variants are available for proteins from any species, extending the framework beyond the human phosphoproteome. Scop3P-Toolkit supports structure-centric analyses including interpretation of PTMs and disease-associated variants, analysis of residue interaction networks and their rewiring across alternative conformations, structural localisation of peptides, and exploration of protein–protein, protein–ligand, and host–pathogen interfaces. Interactive visualisation links sequence annotations, three-dimensional structures, residue interaction networks, and biophysical profiles, enabling coordinated exploration across multiple molecular representations. The toolkit is distributed as Jupyter notebooks, browser-based Voilà applications, and a Galaxy interactive tool, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers. By integrating biological annotation resources into executable, structure-aware workflows, Scop3P-Toolkit enables reproducible interpretation of PTMs, mutations, and proteomics data.
Motivation AlphaFold-based structure prediction has transformed structural biology by enabling accurate protein modelling and providing a powerful framework for inferring protein-protein interactions (PPIs). However, discovering candidate PPIs directly from genome sequences remains a fragmented and largely trial-and-error process, typically requiring separate tools for open reading frame (ORF) prediction, functional annotation, candidate selection, iterative testing of potential partners, manual preparation of individual structural-prediction jobs, and downstream interpretation of confidence metrics. Results We present Protein-Protein Interaction Genomic Finder (ppigFinder), a standalone, cross-platform desktop application that integrates these steps into a project-oriented graphical workflow for genome-based PPI discovery from nucleotide sequence data. ppigFinder combines ORF prediction, functional annotation, genomic-neighbourhood inspection, AlphaFold 3 job generation, remote job submission, and structural-confidence analysis within a single environment. As a proof of concept, we performed a VirD4-centered AlphaFold 3 interactome screen in Xanthomonas citri pv. citri strain 306, modelling VirD4 (ORF2601) against all 4,303 predicted chromosomal ORFs. Ranking by the minimum interchain predicted aligned error (PAE_min) placed all 14 XVIPCD-containing effector candidates within the top 1% of predictions, with the six top-ranked models corresponding to XVIP candidates. The screen also recovered an XVIPCD-containing protein absent from the reference genome annotation and identified high-confidence candidates predicted to bind VirD4 at a surface opposite to the XVIPCD-binding site. Availability and implementation ppigFinder is implemented in Python 3.11 and is freely available under the MIT licence at https://github.com/leepusp/ppigfinder, with documentation and installation instructions for Linux, macOS and Windows. The version described here is archived at [DOI Zenodo — XXXX].
G. U. Oka, Camilla Adan, Celso Vítor Alves Queiroz Calomeno et al.· bioRxiv· 0 citations
Accurate annotation of RNA base-pairing interactions is essential for structural analysis, benchmarking, and data-driven RNA structure prediction. Several tools can extract RNA interactions from three-dimensional coordinates, but their outputs are heterogeneous and may disagree, particularly for non-canonical base pairs. We present EXTRARNAS, a Java-based framework for automated, reproducible, and user-friendly large-scale extraction of RNA structural annotations with multiple tools. EXTRARNAS processes batches of RNA structures specified by PDB identifier and chain, or provided as local PDB files, executes annotation tools through a Docker-based environment, and parses tool-specific outputs using ANTLR4-based grammars. For each structure–tool pair, the framework generates standard BPSEQ files for canonical cis Watson–Crick interactions and introduces BPSEQE, a standardized text format for representing the extended secondary structure, preserving canonical, non-canonical, and multiple interactions per nucleotide. The current prototype supports RNAView, MC-Annotate, and RNAPolis Annotator. We demonstrate EXTRARNAS on eight RNA structures containing triple-helix motifs, comparing extracted canonical pairs against curated BPSEQ references and evaluating the recovery of manually validated Hoogsteen interactions. The results show consistent differences among tools, especially for non-canonical interactions, highlighting the need for standardized representations such as BPSEQE to support reproducible comparison and future consensus-based annotation.
Federico Di Petta, Piermichele Rosati, Piero Hierro Canchari et al.· bioRxiv· 0 citations
Understanding protein mechanisms in health and disease requires characterizing the functional roles of individual amino acid residues. To explore the role of residues and their mutations, we have developed Atlantis, a database that integrates structural and functional information at the human proteome residue level. A graph database enables complex queries and the retrieval of integrated information for multiple functional analysis of protein systems. A Model Context Protocol (MCP) connector allows the interrogation of the resource through Large Language Models (LLMs) or agentic frameworks for biomedical research. Atlantis annotates over 11M residues across 20k human proteins, identifying hundreds thousands intra- and inter-protein contacts in PDB as well as AlphaFoldDB structures. We also provide the possibility to analyze and integrate predicted 3D complexes inputted by the user, and we showcased these features on hundreds of AlphaFold-multimer complexes of GPCRs and LRRK2 interaction networks. The tool is freely accessible at https://atlantis.bioinfolab.sns.it/. GRAPHICAL ABSTRACT
Natalia De Oliveira Rosa, Piergiorgio Ferronato, M. Varisco et al.· bioRxiv· 0 citations
WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches, predicting protein functions from AlphaFold structures using network-based structural homology and filling metabolic model gaps by mapping 75-100% of orphan reactions.
Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.
Diya Mathai, S. Schulze· Journal of Proteome Research· 0 citations
Summary Characterizing proteome complexity in disease contexts is essential for understanding molecular mechanisms and advancing therapeutic development. Mass spectrometry (MS)-based top-down and middle-down proteomics (TDP/MDP) can resolve intact proteoforms — protein molecules carrying a unique combination of isoform sequence and post-translational modifications (PTMs); however, their technical complexity and modest throughput present challenges for experimental planning and limit their broader application. Here, we present ProteoformTracker, an online web tool that prospectively models MS signal and evaluates the feasibility of using TDP/MDP to distinguish a target proteoform from related isoforms and the background proteome. ProteoformTracker takes as input a gene’s annotated isoforms, a novel long-read/assembled transcript, or an rMATS alternative-splicing event, with or without user-specified PTMs, and predicts each proteoform’s MS1 charge-state envelope and exact isotope pattern, scores per-bond MS2 fragmentation propensity, and searches the full reference human proteome for confounding proteins that could share the target’s intact mass or a charge-state m/z peak. ProteoformTracker also supports middle-down workflows via simulated partial protease digestion. Results are rendered as interactive, zoomable MS1 and MS2 visualizations with live resolvability and fragment-ion statistics, letting users incorporate outside evidence into which confounders they compare against. We envision ProteoformTracker as a useful tool for users to plan TDP/MDP experiments targeting specific proteoforms. Availability and implementation ProteoformTracker is implemented in R (Shiny) with a Python backend for exact mass and isotope-pattern calculation and is freely available at http://www.proteoformtracker.org together with the documentation and a walkthrough. The source code is available at https://github.com/HuangLabAtUAB/ProteoformTracker under an MIT license. Supplementary information Supplementary data are available.
Araf Mahmud, Zhi-Hao Zhang, Si Wu et al.· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.