Skip to content
Open access

A Comprehensive Evaluation of Protein Structure Prediction Models for Short Peptides

Jul 2026 · bioRxiv · 1 citation · 37 references
Biology

TL;DR

A comprehensive benchmarking of five state-of-the-art protein structure prediction models demonstrates that prediction accuracy systematically improves with peptide length, and demonstrates that a multi-model consensus approach provides a rational framework for identifying robust structural hypotheses in the absence of experimental reference structures.

Abstract

Short peptides pose distinct challenges for computational structural biology due to their lack of stable tertiary structures, high conformational flexibility, and limited evolutionary signals. To address how modern deep-learning architectures navigate these challenges, we conducted a comprehensive benchmarking of five state-of-the-art protein structure prediction models: AlphaFold2, RoseTTAFold2, ESMFold, OmegaFold, and DMPfold2. Using a curated dataset of experimentally determined short peptide structures (10–49 amino acids) from the Protein Data Bank, we systematically evaluated predictive performance across varying sequence lengths and secondary structure classes. Our results demonstrate that prediction accuracy systematically improves with peptide length. Furthermore, all models perform significantly better on α-helical and mixed-structure peptides compared to β-sheet-rich and intrinsically disordered sequences. Among the evaluated methods, AlphaFold2 and the single-sequence language models, ESMFold and Omegafold proved to be the most consistent and accurate overall. We also observed that internal model confidence scores are imperfectly calibrated for short peptides, necessitating cautious interpretation. Finally, by extending our analysis to the dbAMP3 dataset of uncharacterized antimicrobial peptides, we demonstrate that a multi-model consensus approach provides a rational framework for identifying robust structural hypotheses in the absence of experimental reference structures.

Read PDF

Similar papers

#protein folding Open access Aug 2026

Benchmarking confidence estimation and rescoring for cyclic peptide–protein complex predictions

Cyclic peptides are a rapidly expanding class of therapeutics, but the reliability of deep-learning structure prediction for cyclic peptide–protein complexes has not been systematically evaluated. We assembled a curated benchmark of 111 non-redundant complexes spanning five cyclization chemistries and assessed two co-folding models, Boltz and Protenix, each generating 100 poses per target (22,200 total). Stratifying all poses by complex attributes, we found that disulfide-cyclized peptides and small protein targets (200 or fewer target residues) were predicted significantly worse by both tools, with target size the largest and most consistent effect; overall accuracy nevertheless remained high (median top-pose DockQ of about 0.89, 96–98% of targets Acceptable or better), indicating that pose generation is rarely the bottleneck. Conversely, native model ranking scores correlated only moderately with pose quality (Spearman rank correlations of 0.53–0.66): approximately 12% of poses showed high model ranking score/confidence despite poor pose DockQ quality, and the highest-quality pose was not ranked first for nearly every target. We therefore augmented the native score with externally computed interface descriptors normalized by chain length, principally the per-residue density of inter-chain hydrogen bonds, in a gradient-boosted rescoring model evaluated under target-grouped cross-validation that prevents leakage, improving out-of-fold ROC-AUC for both tools, significantly so for Protenix. Together, these findings identify pose ranking, rather than pose generation, as the major limitation of current cyclic peptide–protein complex prediction and demonstrate that complementary structural features can improve confidence-based pose selection.

Zhe Li, Ye Yuan, Kaiqiang Hu et al. · 0 citations
Open access Jul 2026

A comprehensive dataset of 32 million pentapeptide structures for high-throughput virtual screening.

Small peptides are widely used as binders, modulators, and structural motifs, but their conformational flexibility complicates structure-based analysis and high-throughput screening. We present an open dataset of three-dimensional structures for the complete space of canonical amino-acid pentapeptides: 3,200,000 unique sequences with up to 10 conformers per sequence, for a total of 32,000,000 peptide conformers. Structures were generated directly from sequence using an automated workflow built on UCSF ChimeraX for model construction, Reduce for hydrogen placement, and RDKit for conformer generation and optimization. The dataset is distributed as compressed archives with an accompanying index that maps each sequence and conformer identifier to its coordinate record, enabling efficient download, subset selection, and programmatic access. Technical validation includes symmetry-aware inter-conformer RMSD analysis, Ramachandran quality assessment, and benchmarking against experimentally observed pentapeptide fragments from the Protein Data Bank. Although we focus here on pentapeptides to enable exhaustive sequence coverage, the publicly released workflow is solely based on open-source software and can be applied to other short peptides to generate comparable conformer libraries. This resource supports virtual screening with pre-generated peptide conformer ensembles, method benchmarking, and machine-learning applications in peptide design and protein engineering by removing the need for researchers to repeatedly generate large conformer ensembles from scratch.

Josep-Ramon Codina, E. Dikici, S. Deo et al. · 0 citations
Open access Jul 2026

The accuracy of electrostatic interactions captured by AI protein structure prediction models.

A variant of the U1A protein containing four substitutions to ionizable residues was generated serendipitously due to a miscommunication. Biophysical measurements reveal this variant has twice the helical structure of wild-type U1A and is trimeric, unlike the monomeric wild type. In sharp contrast, structures predicted by deep-learning (AlphaFold2, RoseTTAFold2) and transformer-based tools (OmegaFold, ESMFold) are nearly identical to the wild-type (backbone RMSD < 1 Å). Surprisingly, these models predict ionizable residues buried within the nonpolar core, contradicting established physico-chemical principles. To explore this effect further, we generated sequences containing up to all twelve residues that make up the nonpolar core of U1A. Across thousands of sequences, and depending on the AI model used, the majority of predicted structures contained fully buried ionizable residues while still maintaining the overall U1A fold. We then examined two additional proteins of comparable size, acylphosphatase and the de novo designed TOP7 fold, and observed the same phenomenon: AI models frequently predicted structures with buried ionizable residues that nevertheless retained the parent fold. However, short (50 ns) molecular dynamics simulations with physics-based force fields (CHARMM/AMBER) rapidly relaxed these structures, exposing the ionizable residues. We conclude that while AI-based tools perform exceptionally on natural sequences, they do not reliably encode the physico-chemical principles governing ionizable residue placement. We propose including brief molecular dynamics simulations as a vital validation step for AI-generated structures.

G. Makhatadze · 0 citations
Open access Jul 2026

Mavchen-1: A Conformational Ensemble Platform for Protein–Ligand Pose Prediction That Substantially Outperforms Static Structure Prediction in a Category-Stratified Benchmark

A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.

Ryan Varghese, Pooja Tiwary, Krishil Oswal · 0 citations
Preprint Jul 2026

Accurate structural modeling of chemically diverse molecular interfaces with Vilya-2

Vilya-2 is the structure-prediction oracle that de novo peptide design pipelines require--establishing the all-atom approach as a general foundation for the design and evaluation of de novo peptide therapeutics.

Vilya Research Pascal Sturmfels, Naozumi Hiranuma, M. Salem et al. · 0 citations