Skip to content
Open access

Computational Redesign of an Antifreeze Protein Using Deep Learning

Jun 2026 · bioRxiv · 0 citations · 140 references
Biology

TL;DR

Deep learning-based protein design methods are used to redesign the globular fish antifreeze protein AFPIII, keeping the previously reported ice-binding residues fixed, highlighting the value of deep learning-based protein design methods both for generating AFP variants with desirable properties and for uncovering gaps in existing knowledge of well-characterized AFPs.

Abstract

Antifreeze proteins (AFPs) found in various cold-adapted organisms inhibit ice growth and are of interest for applications in food products, cryopreservation, agriculture, and materials science. Although high-resolution structures are available for several AFPs, the amino acids required for full antifreeze activity remain incompletely defined, and the development of AFP variants with properties such as enhanced solubility, high expression yield, and improved thermostability may further facilitate applications. Here, we used the deep learning model ProteinMPNN to redesign the globular fish antifreeze protein AFPIII, keeping the previously reported ice-binding residues fixed. We readily obtained sequences confidently predicted to adopt AFPIII’s structure and we selected five designed variants for expression, all of which expressed efficiently in E. coli. Circular dichroism spectroscopy showed that two of these variants retained secondary structure elements consistent with AFPIII, whereas the other three exhibited structural differences. One design was predicted and experimentally confirmed to have increased thermostability. All five variants displayed measurable thermal hysteresis activity. However, none reached the activity of wild-type AFPIII, suggesting that maintaining the currently established set of ice-binding residues is not sufficient to fully preserve this AFP’s function; other, unidentified residues can also impact its activity. Our findings highlight the value of deep learning-based protein design methods both for generating AFP variants with desirable properties and for uncovering gaps in existing knowledge of well-characterized AFPs.

Read PDF

Similar papers

Aug 2026

Machine-Learning-Guided Design of Antifreezing Peptides

An unsupervised machine-learning framework that leverages hybrid high-dimensional peptide representations to discover high-performance AFPT families without requiring 3D structures or large labeled data sets is presented and demonstrates how unsupervised hybrid-feature learning can reveal actionable biophysical design rules from sequence data alone.

Nazmul Shuzan, Jialun Wei, Jie Zheng · 0 citations
Open access Jul 2026

ThermoFusion: A Multimodal Deep Learning Framework for Generalizable Prediction of Enzyme Thermostability

Protein thermostability is a critical property for both industrial and biomedical enzyme applications, yet experimental evaluation of mutation-induced stability changes remains laborious and costly. Here, we present ThermoFusion, a hybrid deep learning framework that integrates 3D protein structure embeddings from ThermoMPNN with sequence-based embeddings from the pretrained protein language model ESM2 to predict the effects of single-point mutations on protein stability (ΔΔG). ThermoFusion exhibits robust generalization, maintaining high predictive accuracy across out of distribution sequences with low identity to the training set – a scenario where many other machine learning models, including ThermoMPNN and state-of-the-art tools, perform poorly due to reliance on memorization. Benchmarking on a curated enzyme dataset comprising of 105 enzymes and 3144 mutations shows that ThermoFusion reliably identifies stabilizing mutations while accurately predicting stability for enzymes beyond its training set. These results establish ThermoFusion as a powerful tool for rational enzyme design beyond its training set.

Yao Wei, I. Eberini, Fabian Meyer · 0 citations
Open access Aug 2026

Benchmarking Deep Learning Predictions of Mutation-Induced Fold Switching

A systematic NMR-characterized dataset of mutants of the GA/GB model fold-switching system is presented and it is found that this benchmark revealed variable and position-dependent performance across methods, with certain AlphaFold2-based algorithms able to predict mutant effects at individual sites, indicating some understanding of physical effects of residue substitutions.

Nathaniel R. Felbinger, K. Carillo, Yihong Chen et al. · 0 citations
Open access Jul 2026

The accuracy of electrostatic interactions captured by AI protein structure prediction models.

A variant of the U1A protein containing four substitutions to ionizable residues was generated serendipitously due to a miscommunication. Biophysical measurements reveal this variant has twice the helical structure of wild-type U1A and is trimeric, unlike the monomeric wild type. In sharp contrast, structures predicted by deep-learning (AlphaFold2, RoseTTAFold2) and transformer-based tools (OmegaFold, ESMFold) are nearly identical to the wild-type (backbone RMSD < 1 Å). Surprisingly, these models predict ionizable residues buried within the nonpolar core, contradicting established physico-chemical principles. To explore this effect further, we generated sequences containing up to all twelve residues that make up the nonpolar core of U1A. Across thousands of sequences, and depending on the AI model used, the majority of predicted structures contained fully buried ionizable residues while still maintaining the overall U1A fold. We then examined two additional proteins of comparable size, acylphosphatase and the de novo designed TOP7 fold, and observed the same phenomenon: AI models frequently predicted structures with buried ionizable residues that nevertheless retained the parent fold. However, short (50 ns) molecular dynamics simulations with physics-based force fields (CHARMM/AMBER) rapidly relaxed these structures, exposing the ionizable residues. We conclude that while AI-based tools perform exceptionally on natural sequences, they do not reliably encode the physico-chemical principles governing ionizable residue placement. We propose including brief molecular dynamics simulations as a vital validation step for AI-generated structures.

G. Makhatadze · 0 citations
Open access Jul 2026

Accurate ΔTm Prediction Without Protein Structure Inputs for Biomolecular Stability

Predicting protein stability, like changes in melting temperature (ΔTm) caused by mutations, is a critical task in therapeutic protein engineering and drug discovery. This is reflected by a growing solution space, including both AI-based sequence and structure based methods. This paper demonstrates that accurate ΔTm prediction does not require structural input features, but can achieve state-of-the-art results with a careful training design for large sequence-based protein language models. We combine an autoresearch-inspired setup search with controlled ablation studies and show that a well-tuned sequence-only ESM2-650M model [6] outperforms structure-informed methods in our benchmark, achieving the lowest error (MAE/RMSE) and competitive Pearson correlation without pH or structural inputs. We further show that choices such as loss function, pooling strategy, auxiliary supervision, and finetuning regime materially affect performance.

Daniel Siegismund, Mario Wieser, E. Natali et al. · 0 citations
Open access Jul 2026

Machine Learning-Assisted Evolution of Broadly Functional Enzyme Libraries

Results indicate that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope, and suggest that supervised machine learning can help guide the construction of high-value enzyme libraries with expanded catalytic scope.

Ravi G. Lal, Jason Yang, Ziyan Zhang et al. · 0 citations