Skip to content
Preprint

Commutative Algebra Learning for Protein Flexibility Analysis

Jul 2026 · 1 citation · 34 references
Biology

TL;DR

A commutative algebra-based learning framework, termed CAL, for protein B-factor prediction, that achieves robust and consistent performance across diverse datasets and is competitive with existing state-of-the-art methods.

Abstract

Protein flexibility, commonly quantified by B-factors, is closely related to protein structure and function. However, accurate B-factor prediction remains challenging due to the multiscale nature of protein structures and the complexity of atomic interactions. In this work, we propose a commutative algebra-based learning framework, termed CAL, for protein B-factor prediction. Unlike many biomolecular prediction tasks that rely primarily on global structural representations, B-factor prediction requires an accurate characterization of the local geometric environments surrounding individual atoms. To address this challenge, CAL employs commutative algebra theory to construct localized algebraic descriptors at multiple spatial scales. On a benchmark dataset of 364 proteins, CAL improves prediction accuracy by 34.5\% over the classical Gaussian network model (GNM). Extensive experiments demonstrate that CAL achieves robust and consistent performance across diverse datasets and is competitive with existing state-of-the-art methods. Furthermore, by integrating CAL with machine learning, we develop a blind prediction model capable of cross-protein B-factor prediction. Overall, CAL provides an effective, efficient, and mathematically principled framework for protein flexibility prediction and offers a powerful approach for analyzing and predicting localized structural properties in complex biomolecular systems.

View source

Similar papers

Open access Aug 2026

Entanglement-based continuum conformational landscape of proteins

Motivation With the rapid development of AI methods that predict protein structures from sequence, understanding the structure-function relation increasingly depends on quantitative structural descriptors that are both biologically meaningful and scalable to large datasets. Here, we introduce mathematical topology metrics that quantify the entanglement complexity of a tertiary protein structure while respecting uncrossability constraints. Results By employing only three such metrics across all protein structures in the Protein Data Bank, we represent the proteome structural space in a continuous three-dimensional space. Distances within this space capture structural similarity and correlate with functional similarity. We find that the mathematical entanglement based landscape of protein structural space diversifies with the evolutionary expansion of protein function across species. Moreover, this continuous representation reproduces CATH classifications with high accuracy for major structural classes. These results indicate that these metrics efficiently encode structural features linked to protein function and provide a more informative description than conventional metrics. Availability Data used in this study are available in the Protein Data Bank. Details of the machine learning model used can be found in https://github.com/roshitac/CATH_Classification-. Contact Banu.Ozkan@asu.edu, Eleni.Panagiotou@asu.edu Supplementary information Supplementary data are available at Journal Name online.

P. Malatesta, Roshita S. Chandnani, J. Yalim et al. · 0 citations
Jul 2026

Persistent Manifold Learning of Protein Properties

PML is introduced, a novel computational framework that describes a binding interface as a family of multiscale manifolds, and results indicate that much of what determines binding strength is encoded in the shape of the interface itself, and that a single geometric description serves both classes without hand-tailored features.

Xingjian Xu, Zhe Su, Guo-Wei Wei et al. · 0 citations
Open access Jul 2026

ProteinDock: A physics-informed layer to improve protein-protein docking reliability

It is demonstrated that a truncated version of ProteinDock can be used to choose the optimal prediction among outputs from multiple deep learning-based tools, and shown that this strategy is a computationally efficient alternative to increasing the seed quantity for deep-learning predictions.

G. Rajagopal, Søren C. Spina, Joe Bailey et al. · 0 citations
Review 2026

AI-Driven Protein Research: From Prediction to Design.

This mini review traces the evolution of AI-driven methods in protein research, from early residue-contact prediction using coevolutionary information to transformative breakthroughs, the rise of protein language models (PLMs), and the emerging era of generative design and functional modeling.

Guodong Min, Huan Peng · 0 citations
Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Jul 2026

Density-driven support fields for topological stability in protein structures

We model protein structural stability as a continuous scalar quantity defined over molecular geometry, referred to as the support field. Instead of treating stability as a discrete residue annotation or an empirical score, this representation characterizes protein folds through the combined effects of geometric organization, topological persistence, and local density. Based on this idea, we introduce Support Field Neural Representation Learning (SF-NRL), a topology-guided approach that integrates persistent homology(PH), spatial density estimation, and geometric deep learning to infer residue-wise support directly from protein structures. Persistent topological features are incorporated as structural constraints that modulate local support values across the fold, enabling a continuous description of structural reliability. Across diverse protein families, the inferred support field shows consistent agreement with independent indicators of structural stability and highlights low-support regions associated with conformational flexibility and weak structural integration. By embedding protein structures into a continuous stability landscape, SF-NRL provides an interpretable representation that complements structure prediction models and facilitates systematic identification of structural cores, flexible regions, and functionally relevant motifs. These results demonstrate that topology-informed field representations offer a generalizable and practically useful approach for analyzing protein stability and fold organization.

Jianshi Wang, Yukio Ohsawa · 0 citations