Preprint
Jul 2026
Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
It is concluded that mean attention degradation is largely descriptive rather than prescriptive: function tokens contribute through what their hidden states compute, not through the attention they receive -- with implications for interpretability methodology and attention-score-based inference optimisations such as KV-cache eviction.
Sagar Dangal, Manoj Shakya
· 0 citations