It is shown that ESM-2 exhibits reduced attention on disordered regions, yet still encodes meaningful biological signals, and that both the radius of gyration and individual dynamic contact maps, key characteristics of IDPs, can be obtained from the model logits and embeddings.
Abstract
Protein language models (PLMs) such as ESM-2 encode protein sequences as embeddings for downstream tasks. PLMs are trained on a masked learning objective that leverages evolutionary constraints. While interpretability studies of ESM-2 have focused on folded proteins, their behavior on intrinsically disordered proteins (IDPs), which constitute a substantial fraction of the human proteome and are implicated in numerous diseases, remains understudied. Because IDPs experience different types of evolutionary constraints on their amino acid sequences, we hypothesized that PLMs would behave differently on disordered versus folded regions. Here we show that ESM-2 exhibits reduced attention on disordered regions, yet still encodes meaningful biological signals. The model assigns heightened attention to disease-relevant residues even at high levels of disorder. Moreover, we show that both the radius of gyration and individual dynamic contact maps, key characteristics of IDPs, can be obtained from the model logits and embeddings. These findings suggest PLMs capture valuable information relevant to IDP biology despite their bias toward structured residues.
The utility of HA sites for suggesting candidate binding sites and the biological interpretability of PLM representations is explored, demonstrating the biological interpretability of PLM representations and offers a valuable method to prioritize functionally relevant protein residues for targeted biomedical research.
Sophia J. Pribus, Russ B. Altman, Gowri Nayar· bioRxiv· 0 citations
This work introduces an automated and scalable method for interpreting SAE features in ESM-2 by using geometrically inspired features of the protein $\text{C}_\alpha$ backbone, providing a robust method of annotating proteins activated within SAE neurons at a residue level, providing a bridge between mechanistic interp...
S. Setlur, Djordje Mihajlovic, Darrick Lee· 0 citations
It is shown that PLM embeddings encode patterns correlated with biochemical properties and quantify their contribution to predicting protein fitness, and this technique is readily transferable to problem settings beyond protein fitness prediction.
Paulo Yanez Sarmiento, Pia Francesca Rissom, Manuel Pfeuffer et al.· 0 citations
Abstract Motivation Accurate identification of DNA-binding proteins (DBPs) and RNA-binding proteins (RBPs) is critical for elucidating transcriptional and post-transcriptional regulatory mechanisms. However, existing computational approaches often rely on inferred labels or domain-specific annotations, which limit the...
Hanjin Kim, Sung-Gwon Lee, Joo-Seong Oh et al.· Bioinformatics Advances· 0 citations
Applications to thioredoxins, visual opsins, and Tara Oceans environmental diatom cold-shock proteins show that PLMView can move from interpretable residue-level determinants in well-studied protein families to large-scale environmental functional discovery, linking molecular specialization to ecological distribution a...
Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs)....
R. Joeres, Ilya S. Senatorov, A. Kolchina et al.· 0 citations
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.