Skip to content
Preprint

Task- and dataset-specific information in protein language models

Aug 2026 · 0 citations · 54 references
Computer Science Biology

Abstract

Protein language models (PLMs) have transferred the latest advances from natural language processing to computational biology. These models, trained on large corpora of protein sequence data, are widely used to translate amino acid sequences into latent-space embeddings, ready for use in diverse downstream tasks (DTs). By a common consensus, embeddings from the model's last layer are used, and the model's internal behavior remains poorly understood. We analyzed 13 PLMs across 15 DTs from 11 datasets to investigate the informativeness of embeddings created in intermediate PLM layers. We trained probe models on embeddings from each layer, compared their performance, and computed characteristics of the latent spaces they span to estimate the information they contain, and found that the last layers of PLMs rarely contained embeddings that led to the best results on downstream tasks. Furthermore, we identified a connection between DTs and the distribution across PLMs'layers of the relevant information to predict that task. For example, similarity between the pre-training objective and the objective of predicting properties of individual residues leads to a steady increase in understanding of such tasks across the layers of PLMs. On the other hand, for whole-protein tasks, we observe that the dataset, rather than the task itself, defines PLMs'ability to perform well on a DT. Embeddings from shallow layers of PLMs perform better for datasets that contain deep mutational scan (DMS) data, while datasets containing diverse natural proteins find most useful embeddings in the models'deeper layers. Additionally, we discover that the performance of PLMs drops significantly when tasks are introduced for artificial proteins.

View source

Similar papers

Open access Jul 2026

High-resolution dissection of concept acquisition in different families of protein language models

A high-resolution layer-by-layer interpretability analysis of 8 models from the ESM2 and AMPLIFY families on 22 concepts from human proteome annotations found that these models encode concepts of increasing levels of complexity along their depth: basic physicochemical properties and linear motifs are best captured by early-layer embeddings, secondary structure from subsequent layers, and domain-level semantics from middle layers.

Shawn T. Whitfield, Tom Marty, Robert M. Vernon et al. · 0 citations
Open access Aug 2026

Protein language models and the long tail of functional diversity

It is found that many GigaRef singletons belong to a cluster under alternative parameter settings, suggesting that genomic and metagenomic datasets may require dataset-specific clustering configurations, and it is shown that singletons share mutual information with clustered sequences, making them learnable by PLMs and useful for training.

R. Vinod, Samir Char, Ava A. Amini et al. · 0 citations
Review Open access Jul 2026

Decoding viral protein sequences by large language models

This mini-review summarizes recent developments in devising and applying protein language models for biological sequences, emphasizing viral protein analysis, and outlines a road map for the potential application of LLMs in empowering virology research and pathogen surveillance.

Tianyi Fei, Siqi Li, Ziyue Yang et al. · 0 citations
Jul 2026

A data-centric analysis for efficient semantic knowledge acquisition in word embeddings

A data-centric analysis of semantic knowledge acquisition in word embeddings, focusing on word analogy and semantic similarity shows that, for relational semantics, training-data quality outweighs quantity, and that simple proxy models remain a practical, interpretable tool for efficient data selection.

Aishwarya Jadhav, Mark Anderson, José Camacho-Collados et al. · 0 citations
Open access Jul 2026

Predictions of protein–protein interactions: Learning sequences and structures

A neural network-based pipeline that integrates amino acid sequences with structural features is developed and provides a modular prototype for follow-up, more extensive protein modeling, including larger proteins and sequence of variable sizes.

Carl David Jasper Causin, M. Fyta · 0 citations
Open access Jul 2026

TEDlm: domain-centric protein language models with optional structural pre-training

TEDlm variants also substantially improve zero-shot Molecular Function prediction over ESM2, while matching it on various biophysical property tasks, indicating that signals are largely domain-intrinsic.

Tiejun Wei, S. Kandathil, Daniel W. A. Buchan et al. · 0 citations