Skip to content
Open access

Benchmarking protein sequence and structure search methods for remote homology detection.

Jul 2026 · Genome Biology · 1 citation
Medicine

TL;DR

This benchmark establishes a standardized framework for evaluating protein similarity search methods, providing a practical resource for method selection and a foundation for the development of next-generation approaches capable of addressing diverse homology search challenges.

Abstract

Background

Protein sequence and structure similarity-based search is an important task, which underpins protein annotation, evolutionary analysis, large-scale functional inference, and the exploration of the protein "dark space". The rapid growth of sequence and predicted structure databases has spurred diverse search methods, yet their evaluation remains limited to fold-level similarity and inconsistent benchmarking protocols.

Results

We present a comprehensive benchmark for protein sequence and structure search. Using this framework, we evaluate 14 representative methods spanning sequence alignment, structure alignment, and representation-based approaches across multiple biologically relevant scenarios. Our results show pronounced and context-dependent differences among methods. Structure alignment methods excel at detecting fold-level and geometric similarity, while representation-based searching approaches show advantages in capturing functional similarity under low sequence identity and robustness to predicted structures. Notably, all evaluated methods show limited effectiveness on intrinsically disordered proteins.

Conclusions

This benchmark establishes a standardized framework for evaluating protein similarity search methods, providing a practical resource for method selection and a foundation for the development of next-generation approaches capable of addressing diverse homology search challenges.

Read PDF

Similar papers

Book Open access Aug 2026

CAHS: Context-Aware Homology Search

Protein homology search is foundational to bioinformatics: it supports annotation transfer, structure/function inference, and evolutionary analysis over rapidly expanding sequence repositories (e.g., UniProtKB). Profile hidden Markov models (pHMMs), as implemented in HMMER, remain the most widely trusted approach becau...

Swethasree Bhattaram, D. Bhowmik, Ramakrishnan Kannan · 0 citations
Open access Aug 2026

HInt: interaction-based homology discovery through accelerated genome-scale AlphaFold screening

HInt (Homology by Interaction), an accelerated AlphaFold-based framework that enables practical proteome-scale PPI prediction through biologically informed pre-filtering and optimised high-throughput structure modelling, and provides a general framework for uncovering hidden homologues and expands the conceptual landsc...

Quentin Rouger, P. Paillard, Manon Thomet et al. · 0 citations
Open access Aug 2026

Siteomix: a PyMOL plugin for the detection, alignment, and similarity analysis of protein‒ligand binding sites

Siteomix is an integrated plugin for the PyMOL molecular graphics system that automates the detection of binding pockets via the LIGSITE algorithm, visualizes them as discrete point clouds colored by cavity depth, and performs a two-step alignment combining the rigid iterative closest point (ICP) algorithm with differe...

Kira M. Velieva, E. Skorb, S. Shityakov · 0 citations
Open access Jul 2026

WASP: a pipeline for functional annotation prediction based on AlphaFold structural models

WASP highlights how structural homology can systematically discover annotations missed by sequence-based approaches, predicting protein functions from AlphaFold structures using network-based structural homology and filling metabolic model gaps by mapping 75-100% of orphan reactions.

Giorgia Del Missier, Kiyan Shabestary, Rodrigo Ledesma-Amaro · 0 citations
Open access Sep 2026

Foldseek-Interface reveals a protein interface universe far from complete

Foldseek-Interface is presented, a method that converts 3D interface structures into searchable sequences to enable fast alignment and clustering of protein interaction interfaces and matches the accuracy of state-of-the-art tools while running up to 230 times faster.

J. M. Strom, Sooyoung Cha, R. Kim et al. · 1 citation · ⚡1
Open access Aug 2026

ProtPen Combines Sequence- and Structure-based Approaches to Facilitate Protein Function Predictions on a Proteome-wide Scale.

Proteins of unknown function represent a significant gap in our understanding of biological processes, encompassing large portions of the proteomes of many organisms, especially prokaryotes. Addressing this gap is critical to understanding the biology and pathogenicity of such organisms. We introduce ProtPen, an open-s...

Diya Mathai, S. Schulze · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.