Minimal Data Maximal Insight (MDMI), a two-stage structure-guided computational pipeline that designs functional peptide variants using only a small, annotated dataset, demonstrates that structure-informed pipelines can uncover remote functional sequence space from minimal data.
Abstract
Discovering functional peptides across vast sequence space remains a formidable challenge, particularly when experimental training data is scarce. We present Minimal Data Maximal Insight (MDMI), a two-stage structure-guided computational pipeline that designs functional peptide variants using only a small, annotated dataset. Rather than relying on sequence information alone, MDMI integrates three-dimensional structural features derived from predicted peptide-protein complexes into a machine learning model that captures interface geometry and binding energetics. This structure-aware predictor, paired with a genetic algorithm for sequence exploration, reduced false positives from 70% to close to zero in an all-negative benchmark panel compared with a sequence-only model in computational benchmarking, and produced approximately four-fold more high-confidence in silico binders than state-of-the-art peptide/protein design baselines. Using the split-GFP system as a testbed, where fluorescence provides a direct functional readout of peptide-protein complementation, MDMI identified peptides with up to 38% sequence divergence from wild-type in Stage 1 while retaining measurable activity. In Stage 2, motif-guided recombination of successful Stage 1 variants produced highly divergent yet functional peptides bearing over 50% sequence difference from wild-type, revealing two distinct functional clusters in sequence space. As further validation, a top-performing candidate expressed as a full-length GFP fusion retained a GFP-like emission profile, supporting formation of a fluorescent GFP-like scaffold. These results demonstrate that structure-informed pipelines can uncover remote functional sequence space from minimal data, with broad implications for peptide and therapeutic analog discovery.
A pipeline reformulating kinase-substrate modeling as a Bayesian inference problem is presented and it is revealed that the interaction types and distances to the catalytic pocket significantly influence pathogenicity scores.
Jinyuan Hu, Shimian Li, Yue Xue et al.· Journal of Chemical Informat...· 0 citations
Accurate computational prediction of enzyme function, standardized by Enzyme Commission (EC) numbers, is essential for large-scale genome annotation and generative enzyme design. However, it remains unclear whether state-of-the-art predictors learn the intrinsic structural determinants of catalytic activity or merely rely on global sequence similarity to annotated homologues. To address this gap, we introduce EnzymARC, a novel benchmark dataset of putative non-functional decoy sequences generated via structure-guided, systematic disruption of active sites (targeting catalytic residues and surrounding 5 Å, 10 Å, and 15 Å radii) from experimentally annotated enzymes. We evaluated three distinct prediction paradigms against this dataset: homology-based annotation (DIAMOND), contrastive learning with protein language models (CLEAN), and a deep learning model incorporating non-enzyme discrimination (DeepEC). Our findings reveal that current models are highly vulnerable to phylogenetic shortcuts. Both DIAMOND and CLEAN exhibited false positive rates exceeding 90% for low-perturbation decoys, confidently assigning the original EC numbers despite the destruction of the catalytic machinery. While DeepEC demonstrated improved sensitivity at higher perturbation levels—highlighting the benefit of negative training examples—all models struggled to identify targeted active-site disruptions. We demonstrate that modern EC predictors largely fail to distinguish catalytically incompetent variants from functional enzymes, and we propose that integrating structure-aware negative examples into both training and benchmarking is critical for developing functionally robust models in computational enzymology.
João Sartori, Ana Carolina Ramos Guimarães, Lucas de Almeida Machado· bioRxiv· 0 citations
The results demonstrate the effectiveness of integrating multi-scale and multi-modal representations with cross-scale alignment for protein–RNA affinity prediction, and suggest that M2-PRNet can highlight relevant RNA-binding regions and support preliminary discrimination between strong and weak binders when plausible complex structures are available.
Junkai Wang, G. Luo, Yun-Song Yang et al.· Bioinformatics· 0 citations
A category-stratified, statistically powered benchmark comparing pose prediction from receptor conformational ensembles against AlphaFold2, used as a matched static-structure baseline, across 29 protein–ligand systems spanning cryptic-pocket, induced-fit, water-mediated, and autoimmune-indication target classes is presented.
Ryan Varghese, Pooja Tiwary, Krishil Oswal· bioRxiv· 0 citations
Protein-peptide interactions are central to cellular signaling and to a growing class of peptide therapeutics, yet the datasets used to develop and benchmark computational methods for protein-peptide modeling remain poorly standardized. Available databases prioritize comprehensive coverage but require task-specific curation, while published benchmarks are typically distributed as static collections built with heterogeneous curation, quality-filtering, redundancy-reduction, and sampling strategies, limiting reproducibility and cross-study comparison. We present PepXPro, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria. PepXPro is organized into three components: Scrape, for deterministic curation of protein-peptide complex entries from public resources; GenSample, for constructing configurable subsets under explicit quality, redundancy, and sampling constraints; and Benchmark, for evaluating candidate subsets and selecting a nonredundant, representative, general-purpose benchmark for distribution. Starting from PDBbind and complementary resources, the curation pipeline yields a pool of proteinpeptide complex entries that retains chemically complex cases, including disulfide- linked cyclic peptides, which are commonly excluded from existing benchmarks. We release PepXPro Benchmark v1, a benchmark comprising 70 non-redundant protein- peptide complexes with experimentally determined structures and binding affinities. The underlying framework provides an extensible foundation for reproducible protein- peptide benchmark construction.
Affinity reagents such as antibodies are indispensable for interrogating proteins’ biological function. Yet they are costly and frequently unreliable, with unknown sequences, posing challenges to reproducible experimental research. Deep learning-based protein design can now in silico generate affinity reagents achieving reliable experimental success rates, but has remained largely confined to specialist laboratories. Here we present the Human Bindome, a proteome-scale atlas of high-confidence in silico protein binder candidates. By embedding the experimentally benchmarked BindCraft method in an accelerated, parallelized framework with automated domain-level target selection, we generated 306,146 binder candidates covering 8,296 human proteins (40.9% of the full proteome). Every candidate carries a defined sequence, a predicted binder-target structure model, and in silico confidence metrics. We characterize proteome-wide coverage and show that binder epitopes frequently overlap functional sites. This positions the Bindome as a resource of genetically encodable perturbagens for site-specific, modular control of protein function. The Bindome is freely available through a web interface (https://bindome.epfl.ch), with agentic, natural-language querying and as data splits for machine-learning model development. We anticipate that the Bindome will be valuable for the scientific community by providing affinity and perturbation reagents with broad applications in dissecting biological mechanisms as well as in drug and target discovery.
Julius Wenckstern, Anna M. Díaz-Rovira, Julia A. Kuhn et al.· bioRxiv· 0 citations