Hyper-Fold is introduced, a rank-K separable convolutional backbone approaching this ceiling at message-passing cost, suggesting that a sufficiently expressive 3D backbone recovers information that fusion architectures previously borrowed from evolution-scale pretraining.
Abstract
Protein structure modeling rests on a single computational primitive: the interaction between what a residue is (sequence content) and where it sits (three-dimensional geometry). What is the expressive limit of this layer class? We show that the complete bilinear operator over content-geometry outer products--the sufficient statistic of all second-order interactions--is the expressive ceiling, while the additive message passing of mainstream geometric GNNs is provably blind to content-geometry binding. We then introduce Hyper-Fold, a rank-K separable convolutional backbone approaching this ceiling at message-passing cost: each radius neighborhood is organized into a sequence hyperedge and a contact hyperedge, modulated by an edge-conditioned matrix-valued operator factorized into K learned basis operators with geometry-generated coefficients. Across enzyme function prediction, fold classification, and ligand binding site detection, Hyper-Fold and its hierarchical variant Hyper-Fold-Deep achieve the best results among protein-specific structure encoders; Hyper-Fold-Pocket, an anchored set-prediction head, surpasses UniSite-3D on UniSite-DS and two zero-shot benchmarks with no sequence language model features, 68x fewer parameters, and 4.8x lower latency--suggesting that a sufficiently expressive 3D backbone recovers information that fusion architectures previously borrowed from evolution-scale pretraining.
OmniScore is introduced, a universal structure-based framework that learns a shared geometry-aware representation of complexes once and then adapts it to downstream scoring through lightweight task-specific heads, suggesting that geometry-aware pretraining can provide a reusable scoring backbone for tasks that depend o...
Predicting ligand binding sites on protein surfaces requires capturing complex local geometries and satisfying physical constraints. Existing voxel-based methods suffer from high computational costs and rotation sensitivity, while standard point-cloud GNNs often lack geometric completeness—failing to distinguish chiral...
Kang-Xin Chen, Jie-Yu Zhao, Jin-Li Hu et al.· Proceedings of the Thirty-Fi...· 0 citations
Design of functional biological sequences such as DNA, RNA and peptides has wide-ranging applications in nanomaterials, bio-sensing and medicine. One common challenge across applications is the need to optimize complex high-dimensional properties such as target emission spectra of DNA-mediated fluorescent nanoclusters,...
Elham Sadeghi, I-Hsin Lin, Xian-Qi Deng et al.· Proceedings of the 32nd ACM...· 0 citations
A cheap density prior is introduced over natural protein activations and keeps only the candidates that remain typical under it, a training-free post-hoc step the authors call Mahalanobis filtering that improves both the property score and the structural plausibility of the sequences it keeps at negligible cost, withou...
Shuibai Zhang, Xin-Chi Liu, Fred Zhangzhi Peng et al.· 0 citations
Inverse FoldDir is a structure-conditioned protein redesign method that combines structural recovery, user control, experimental validation, and a natural route toward future property-guided sampling that performs iterative denoising on the amino acid probability simplex.
Alp Tartici, M. Stojkovic, An-Ru Tian et al.· bioRxiv· 0 citations
Accurate identification of near-native ligand binding poses is a central challenge in structure-based drug design. From a physical point of view, the successful construction of a protein-ligand complex structure is dependent on whether protein and ligand can form enough atomically pairwise interactions that result in a...
Zhen-Qiang Zhang, Zhihao Wang, Yang Liu et al.· Journal of Chemical Informat...· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.