HInt (Homology by Interaction), an accelerated AlphaFold-based framework that enables practical proteome-scale PPI prediction through biologically informed pre-filtering and optimised high-throughput structure modelling, and provides a general framework for uncovering hidden homologues and expands the conceptual landscape of protein homology inference.
Abstract
Identifying homologous proteins across deep evolutionary distances remains a major challenge because sequence and structural similarity progressively become undetectable over time. Although protein-protein interactions (PPIs) are often constrained by function and evolution, whether conserved interaction interfaces can provide an independent signal for homology detection has remained largely unexplored owing to the computational cost of proteome-scale interaction prediction. Here we introduce HInt (Homology by Interaction), an accelerated AlphaFold-based framework that enables practical proteome-scale PPI prediction through biologically informed pre-filtering and optimised high-throughput structure modelling. Using HInt, we establish interaction-based similarity as a third axis of homology detection. We show that conserved interaction interfaces reveal homologous relationships that remain inaccessible to conventional sequence- and structure-based approaches. Application of HInt to both prokaryotic and eukaryotic systems, together with experimental validation, uncovered a previously unrecognised VirB5 pilus-tip protein in the F-plasmid type IV secretion system and a previously unannotated F-box-like protein in the Saccharomyces cerevisiae ubiquitin-proteasome system. By enabling practical proteome-scale interaction screening, HInt provides a general framework for uncovering hidden homologues and expands the conceptual landscape of protein homology inference.
Protein-protein interactions mediate a vast range of cellular functions, requiring diverse modes of binding. While recent years have seen major efforts to chart and classify the protein structure universe, we lack comparable methods to assess and cluster that diversity in interface structure at interactome scale. Here, we present Foldseek-Interface, a method that converts 3D interface structures into searchable sequences to enable fast alignment and clustering of protein interaction interfaces. It matches the accuracy of state-of-the-art tools while running up to 230 times faster. Applying it to all biological assemblies in the PDB, we cluster 3.1 million dimers into 77,167 interface clusters and use this resource to characterise interface diversity, evolution, and pathogen mimicry. Application of Foldseek-Interface to resources of predicted protein complex structures rapidly revealed putatively novel interface types worth further experimental interrogation. Foldseek-Interface and the interface cluster resource are freely available as webservers for search (https://search.foldseek.com/interface) and exploration (https://interface.foldseek.com). Contact: martin.steinegger@snu.ac.kr, k.luck@imb-mainz.de
J. M. Strom, Sooyoung Cha, R. Kim et al.· bioRxiv· 0 citations
Whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins and suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation are suggested.
Lino Riepenhausen, Francesco Costa, Antonina Andreeva et al.· bioRxiv· 0 citations
WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches, predicting protein functions from AlphaFold structures using network-based structural homology and filling metabolic model gaps by mapping 75-100% of orphan reactions.
Motivation AlphaFold-based structure prediction has transformed structural biology by enabling accurate protein modelling and providing a powerful framework for inferring protein-protein interactions (PPIs). However, discovering candidate PPIs directly from genome sequences remains a fragmented and largely trial-and-error process, typically requiring separate tools for open reading frame (ORF) prediction, functional annotation, candidate selection, iterative testing of potential partners, manual preparation of individual structural-prediction jobs, and downstream interpretation of confidence metrics. Results We present Protein-Protein Interaction Genomic Finder (ppigFinder), a standalone, cross-platform desktop application that integrates these steps into a project-oriented graphical workflow for genome-based PPI discovery from nucleotide sequence data. ppigFinder combines ORF prediction, functional annotation, genomic-neighbourhood inspection, AlphaFold 3 job generation, remote job submission, and structural-confidence analysis within a single environment. As a proof of concept, we performed a VirD4-centered AlphaFold 3 interactome screen in Xanthomonas citri pv. citri strain 306, modelling VirD4 (ORF2601) against all 4,303 predicted chromosomal ORFs. Ranking by the minimum interchain predicted aligned error (PAE_min) placed all 14 XVIPCD-containing effector candidates within the top 1% of predictions, with the six top-ranked models corresponding to XVIP candidates. The screen also recovered an XVIPCD-containing protein absent from the reference genome annotation and identified high-confidence candidates predicted to bind VirD4 at a surface opposite to the XVIPCD-binding site. Availability and implementation ppigFinder is implemented in Python 3.11 and is freely available under the MIT licence at https://github.com/leepusp/ppigfinder, with documentation and installation instructions for Linux, macOS and Windows. The version described here is archived at [DOI Zenodo — XXXX].
G. U. Oka, Camilla Adan, Celso Vítor Alves Queiroz Calomeno et al.· bioRxiv· 0 citations
Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-source pipeline that facilitates protein function prediction by combining eggNOG-mapper for sequence-based annotation with Foldseek for rapid structural similarity searches using AlphaFold-predicted protein structures. Annotation results from both tools are merged and enriched with UniProt metadata to produce a comprehensive output suitable for downstream analysis. The pipeline requires only a FASTA input file with UniProt identifiers, and is designed to analyze data sets on the scale of whole proteomes. Benchmarking on a curated data set of well-characterized Pseudomonas aeruginosa proteins demonstrated an annotation accuracy of >90%, and highlighted the complementarity of sequence- and structure-based methods. Further evaluation of ProtPen included its application to biologically relevant data sets, comprising proteins of unknown function that exhibited significant differential abundances in a proteomics data set of P. aeruginosa, and uncharacterized glycoproteins from Haloferax volcanii. ProtPen is readily extensible to incorporate additional protein function prediction tools. In summary, this pipeline facilitates the systemwide annotation of proteins of unknown function from proteomic data sets and whole proteomes.
Diya Mathai, S. Schulze· Journal of Proteome Research· 0 citations
Improvements in computational protein structure prediction have enabled searches for remote homologs of proteins whose molecular function remain unknown. However, the reliability of functional annotations from such searches has not been systematically quantified. In this study, we assembled a time-split benchmark from Pfam families annotated as domains of unknown function (DUFs) in Pfam 28.0 and were subsequently assigned a function in the current release (Pfam 38.0). These retrospective-DUFs provide ground truth for assessing functional annotation from remote homology searches. In our benchmark analysis, we paired retrospective-DUFs with a difficulty-matched arm from known domains to distinguish query difficulty from the method’s performance. Across four model proteomes (yeast, C. elegans, Drosophila, and mice), Foldseek searches against AlphaFold/Swiss-Prot, PDB100, and CATH50 recovered the later-assigned function for 15.9% of retrospective DUF queries, compared with 30.1% of matched known-domain queries, after masking for self-family and self-clan level hits to remove circularity. At a fixed confidence cut-off (qTM ≥ 0.5), 55.4% of informative calls on retrospective DUFs were incorrect, establishing an error model for prospective use. Furthermore, in our comparison of tools for homology searches, Foldseek outperformed MMseq2-based sequence search but was not statistically separable from the ESM-2 protein language model embedding baseline. Applying the calibrated structural homology search pipeline to 296 currently unannotated DUF queries in Pfam 38.0 yielded 50 confident, specific functional assignments. We provide the benchmark and ranked candidates as a resource to facilitate functional studies.
Jodh S. Pannu, Kendall Green, Jeffrey Vedanayagam· bioRxiv· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.