Mechanistic Interpretability of Protein Language Models Reveals Encoded Structural and Functional Properties of Intrinsically Disordered Proteins
It is shown that ESM-2 exhibits reduced attention on disordered regions, yet still encodes meaningful biological signals, and that both the radius of gyration and individual dynamic contact maps, key characteristics of IDPs, can be obtained from the model logits and embeddings.
Lauren Naworski, Lydia L. Good, Rob M. Scrutton et al.
· bioRxiv · 0 citations