This benchmark establishes a standardized framework for evaluating protein similarity search methods, providing a practical resource for method selection and a foundation for the development of next-generation approaches capable of addressing diverse homology search challenges.
Abstract
Background
Protein sequence and structure similarity-based search is an important task, which underpins protein annotation, evolutionary analysis, large-scale functional inference, and the exploration of the protein "dark space". The rapid growth of sequence and predicted structure databases has spurred diverse search methods, yet their evaluation remains limited to fold-level similarity and inconsistent benchmarking protocols.
Results
We present a comprehensive benchmark for protein sequence and structure search. Using this framework, we evaluate 14 representative methods spanning sequence alignment, structure alignment, and representation-based approaches across multiple biologically relevant scenarios. Our results show pronounced and context-dependent differences among methods. Structure alignment methods excel at detecting fold-level and geometric similarity, while representation-based searching approaches show advantages in capturing functional similarity under low sequence identity and robustness to predicted structures. Notably, all evaluated methods show limited effectiveness on intrinsically disordered proteins.
Conclusions
This benchmark establishes a standardized framework for evaluating protein similarity search methods, providing a practical resource for method selection and a foundation for the development of next-generation approaches capable of addressing diverse homology search challenges.
Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach becau...
Swethasree Bhattaram, D. Bhowmik, Ramakrishnan Kannan· Proceedings of the 32nd ACM...· 0 citations
HInt (Homology by Interaction), an accelerated AlphaFold-based framework that enables practical proteome-scale PPI prediction through biologically informed pre-filtering and optimised high-throughput structure modelling, and provides a general framework for uncovering hidden homologues and expands the conceptual landsc...
Quentin Rouger, P. Paillard, Manon Thomet et al.· bioRxiv· 0 citations
Siteomix is an integrated plugin for the PyMOL molecular graphics system that automates the detection of binding pockets via the LIGSITE algorithm, visualizes them as discrete point clouds colored by cavity depth, and performs a two-step alignment combining the rigid iterative closest point (ICP) algorithm with differe...
Kira M. Velieva, E. Skorb, S. Shityakov· Journal of Computer-Aided Mo...· 0 citations
WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches, predicting protein functions from AlphaFold structures using network-based structural homology and filling metabolic model gaps by mapping 75-100% of orphan reactions.
Foldseek-Interface is presented, a method that converts 3D interface structures into searchable sequences to enable fast alignment and clustering of protein interaction interfaces and matches the accuracy of state-of-the-art tools while running up to 230 times faster.
J. M. Strom, Sooyoung Cha, R. Kim et al.· bioRxiv· 1 citation· ⚡1
Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-s...
Diya Mathai, S. Schulze· Journal of Proteome Research· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.