Skip to content
#protein folding Preprint

Hyper-Fold: Exploring the Expressive Limit of Sequence-Geometry Learning for Proteins via Hypergraph Modeling

Aug 2026 · 0 citations · 40 references
Computer Science

TL;DR

Hyper-Fold is introduced, a rank-K separable convolutional backbone approaching this ceiling at message-passing cost, suggesting that a sufficiently expressive 3D backbone recovers information that fusion architectures previously borrowed from evolution-scale pretraining.

Abstract

Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (three-dimensional geometry). What is the expressive limit of this layer class? We show that the complete bilinear operator over content-geometry outer products--the sufficient statistic of all second-order interactions--is the expressive ceiling, while the additive message passing of mainstream geometric GNNs is provably blind to content-geometry binding. We then introduce Hyper-Fold, a rank-K separable convolutional backbone approaching this ceiling at message-passing cost: each radius neighborhood is organized into a sequence hyperedge and a contact hyperedge, modulated by an edge-conditioned matrix-valued operator factorized into K learned basis operators with geometry-generated coefficients. Across enzyme function prediction, fold classification, and ligand binding site detection, Hyper-Fold and its hierarchical variant Hyper-Fold-Deep achieve the best results among protein-specific structure encoders; Hyper-Fold-Pocket, an anchored set-prediction head, surpasses UniSite-3D on UniSite-DS and two zero-shot benchmarks with no sequence language model features, 68x fewer parameters, and 4.8x lower latency--suggesting that a sufficiently expressive 3D backbone recovers information that fusion architectures previously borrowed from evolution-scale pretraining.

View source

Similar papers

Open access Aug 2026

OmniScore: Universal Scoring of Diverse Biomolecular Complexes via Equivariant Geometry-Aware Discrete Representation Learning

OmniScore is introduced, a universal structure-based framework that learns a shared geometry-aware representation of complexes once and then adapts it to downstream scoring through lightweight task-specific heads, suggesting that geometry-aware pretraining can provide a reusable scoring backbone for tasks that depend o...

Tien-Cuong Bui, Junsu Ko, Ju-Yong Lee · 0 citations
Conference Open access Sep 2026

UniPocket: Physics-Aware Geometric Graph Learning with Manifold Completeness for Ligand-Specific Binding Site Prediction

Predicting ligand binding sites on protein surfaces requires capturing complex local geometries and satisfying physical constraints. Existing voxel-based methods suffer from high computational costs and rotation sensitivity, while standard point-cloud GNNs often lack geometric completeness—failing to distinguish chiral...

Kang-Xin Chen, Jie-Yu Zhao, Jin-Li Hu et al. · 0 citations
Book Open access Aug 2026

Geometry-Preserving Supervised Biological Sequence Design

Design of functional biological sequences such as DNA, RNA and peptides has wide-ranging applications in nanomaterials, bio-sensing and medicine. One common challenge across applications is the need to optimize complex high-dimensional properties such as target emission spectra of DNA-mediated fluorescent nanoclusters,...

Elham Sadeghi, I-Hsin Lin, Xian-Qi Deng et al. · 0 citations
#protein folding Preprint Aug 2026

Off-Manifold Collapse in Guided Protein Language Models

A cheap density prior is introduced over natural protein activations and keeps only the candidates that remain typical under it, a training-free post-hoc step the authors call Mahalanobis filtering that improves both the property score and the structural plausibility of the sequences it keeps at negligible cost, withou...

Shuibai Zhang, Xin-Chi Liu, Fred Zhangzhi Peng et al. · 0 citations
#protein folding Open access Sep 2026

Inverse FoldDir: Structure-conditioned Protein Sequence Design by Dirichlet Flow Matching

Inverse FoldDir is a structure-conditioned protein redesign method that combines structural recovery, user control, experimental validation, and a natural route toward future property-guided sampling that performs iterative denoising on the amino acid probability simplex.

Alp Tartici, M. Stojkovic, An-Ru Tian et al. · 0 citations
Jul 2026

Native Contact Ratio as a Topological Metric for Machine Learning Based Molecular Docking.

Accurate identification of near-native ligand binding poses is a central challenge in structure-based drug design. From a physical point of view, the successful construction of a protein-ligand complex structure is dependent on whether protein and ligand can form enough atomically pairwise interactions that result in a...

Zhen-Qiang Zhang, Zhihao Wang, Yang Liu et al. · 0 citations

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.