Skip to content
#protein folding Open access

Mechanistic Interpretability of Protein Language Models Reveals Encoded Structural and Functional Properties of Intrinsically Disordered Proteins

Sep 2026 · bioRxiv · 0 citations · 39 references
Biology

TL;DR

It is shown that ESM-2 exhibits reduced attention on disordered regions, yet still encodes meaningful biological signals, and that both the radius of gyration and individual dynamic contact maps, key characteristics of IDPs, can be obtained from the model logits and embeddings.

Abstract

Protein language models (PLMs) such as ESM-2 encode protein sequences as embeddings for downstream tasks. PLMs are trained on a masked learning objective that leverages evolutionary constraints. While interpretability studies of ESM-2 have focused on folded proteins, their behavior on intrinsically disordered proteins (IDPs), which constitute a substantial fraction of the human proteome and are implicated in numerous diseases, remains understudied. Because IDPs experience different types of evolutionary constraints on their amino acid sequences, we hypothesized that PLMs would behave differently on disordered versus folded regions. Here we show that ESM-2 exhibits reduced attention on disordered regions, yet still encodes meaningful biological signals. The model assigns heightened attention to disease-relevant residues even at high levels of disorder. Moreover, we show that both the radius of gyration and individual dynamic contact maps, key characteristics of IDPs, can be obtained from the model logits and embeddings. These findings suggest PLMs capture valuable information relevant to IDP biology despite their bias toward structured residues.

Read PDF

Similar papers

Open access Aug 2026

Interpreting Protein Language Models: high attention sites predict functional regions

The utility of HA sites for suggesting candidate binding sites and the biological interpretability of PLM representations is explored, demonstrating the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.

Sophia J. Pribus, Russ B. Altman, Gowri Nayar · 0 citations
Preprint Aug 2026

Interpreting Latent Protein Language Model Features with Geometric Annotations

This work introduces an automated and scalable method for interpreting SAE features in ESM-2 by using geometrically inspired features of the protein $\text{C}_\alpha$ backbone, providing a robust method of annotating proteins activated within SAE neurons at a residue level, providing a bridge between mechanistic interp...

S. Setlur, Djordje Mihajlovic, Darrick Lee · 0 citations
Preprint Aug 2026

Interpreting Protein Language Model Embeddings via Orthogonal Projection for Protein Fitness Prediction

It is shown that PLM embeddings encode patterns correlated with biochemical properties and quantify their contribution to predicting protein fitness, and this technique is readily transferable to problem settings beyond protein fitness prediction.

Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer et al. · 0 citations
Open access Aug 2026

Interpretable prediction of nucleic acid-binding proteins using a protein language model

Abstract Motivation Accurate identification of DNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) is critical for elucidating transcriptional and post-transcriptional regulatory mechanisms. However, existing computational approaches often rely on inferred labels or domain-specific annotations, which limit the...

Hanjin Kim, Sung-Gwon Lee, Joo-Seong Oh et al. · 0 citations
Open access Aug 2026

PLMView: collaborative protein language model representations for fast and scalable specialized protein function inference

Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution a...

Vinh-Son Pho, Alessandro Natale Bianchi, Mattéo Scarsini et al. · 0 citations
#machine learning Preprint Aug 2026

Task- and dataset-specific information in protein language models

Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs)....

R. Joeres, Ilya S. Senatorov, A. Kolchina et al. · 0 citations

Related blog posts

Google DeepMind Blog Sep 30, 2026

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

MIT News · Artificial Intelligence Aug 27, 2026

Looking beyond natural sequences

A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.