Skip to content
Preprint

Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data

Aug 2026 · 0 citations · 27 references
Computer Science Mathematics

TL;DR

Treating rows of a transformer's attention matrix as compositional data separates them exactly: the Aitchison distance splits orthogonally into a sink term and a content term, entropy splits by an exact identity, and the content distance is characterized by invariances the transformer itself possesses.

Abstract

Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability lands on a single \emph{sink} token, usually the first. Standard tools for comparing attention rows (cosine similarity, Jensen--Shannon divergence, Shannon entropy) therefore hinge on a choice papers rarely report: keep the sink, or drop it and renormalize. This choice can reverse conclusions. On ten pretrained models from five families, 17--47% of verdicts about which of two heads is more similar flip with the convention, and the most prominent structure in a standard BERT head-clustering pipeline is an artifact of it. The reason is that one-number summaries mix two questions: how much attention the sink takes, and how the rest is divided among the content tokens. Treating rows as compositional data separates them exactly: the Aitchison distance splits orthogonally into a sink term and a content term, entropy splits by an exact identity, and the content distance is characterized by invariances the transformer itself possesses. The separation matters in practice: most measured entropy collapse during training is the sink growing, not attention sharpening (30% of the drop at 70M parameters, 95% at 1B, 79% at 1.4B), and pruning heads with the wrong channel can inflate perplexity more than a hundredfold. We map where each convention is safe, test a frozen out-of-sample predictor (one confirmation, one abstention, one failure), and release code regenerating every number.

View source

Similar papers

Review Aug 2026

Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off

This survey traces attention from Bahdanau-Luong alignment through the Transformer and into vision architectures, and reviews fixed and learned sparse attention, linear attention, IO-aware exact algorithms including FlashAttention, and state-space alternatives including Mamba.

Aditya Singh · 0 citations
#machine learning Preprint Sep 2026

Retrieval Capacity of Self-Attention Under Competition

How many tokens from its context does a language model actually use, and what determines that number? We study this question through self-attention. Without retraining, we retain only the tokens with the highest attention weights at each head, layer, and query, keeping their original weights unchanged. By varying the s...

Timur Mudarisov, M. Burtsev, Radu State · 0 citations
#machine learning Preprint Sep 2026

DimPO: Dimensionality Reduction for Attention using Preference Optimization

A linear projection can reduce the dimension of query and key vectors without updating the pretrained model, but it remains unclear which training objective best preserves model behavior. We ask whether preferences over keys and attention mass on the highest-weighted keys provide a better signal than matching the full...

Vojtěch Lanz, Yu-Fei Cui, Prasanna Parthasarathi · 0 citations
#machine learning Preprint Sep 2026

KuaFu: Compressing Long User Behavior into Understanding at Billion Scale

KuaFu is a unified behavior-compression layer whose minimal unit is one behavior item, with fidelity-oriented four-stage training and layered intermediate evaluation, which matches or exceeds uncompressed single-task production models on all five headline metrics, raises per-GPU throughput by 37%-350%, and saves 190 GP...

Jia-Hao Hui, Lin Zhu, Yi-Sheng Hu et al. · 0 citations
#machine learning Preprint Sep 2026

Latest Exact Match Attention

We introduce latest exact match attention (LEMA), an attention variant for transformers where queries and keys are binarized and each query attends only to the latest exactly matching key. We prove that LEMA transformers with chain of thought can simulate word-RAMs, as was recently shown for the less restrictive rightm...

M. Brösamle · 0 citations
#artificial intelligence Preprint Sep 2026

Rules Amortize, Pairings Don't: Linguistic Structure Determines What Latent Task Representations Can Replace In-Context Learning

In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at zero-shot inference cost, but recent theory shows a static vector acts as a single synthetic demonstration and must fail on high-rank mappings such as word-level bijections....

Gunmay Jhingran · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.