Skip to content

How Linear Attention Remembers

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

Together, these results reveal how fixed-size recurrent memory supports selective recall despite shared storage, while exposing the interference and capacity limits that distinguish it from token-addressable KV memory.

Abstract

Linear attention replaces the growing key--value (KV) cache of standard attention with a fixed-size recurrent state, substantially reducing memory growth with context length. This efficiency, however, changes how past information is stored: many tokens must share and repeatedly update the same memory. We study how this recurrent state functions as a memory system. Using an analytical decomposition together with controlled causal interventions in pretrained GLA and GDN models, we trace how recalled information is written, retained, and later accessed. We find that fact-specific information enters recurrent memory through concentrated, content-dependent writes and is later accessed through concentrated query-time read pathways. Multiple facts can remain selectively accessible within the same state, yet their internal representations exhibit cross-fact causal coupling rather than independent KV-like storage. As memory load increases, both recall and targeted editability degrade, whereas elapsed context alone has a substantially smaller effect within the tested regime. Causal interventions on subsequent writes further show that interference is shaped by their overlap with existing memory. Finally, in hybrid architectures that combine recurrent layers with full attention, the runtime memory directly supporting recall shifts predominantly to the full-attention KV state. Together, these results reveal how fixed-size recurrent memory supports selective recall despite shared storage, while exposing the interference and capacity limits that distinguish it from token-addressable KV memory.

View source

Similar papers

#machine learning Preprint Sep 2026

CyFA: Linear Sequence Modeling with Relative-Time-Partitioned Memory

Linear RNNs offer linear-time sequence processing and constant-memory decoding, but their fixed-size recurrent states must accommodate all past key--value associations. Existing forgetting mechanisms and Delta Rule updates reduce interference by selectively clearing or correcting the state, yet earlier associations can...

Yi-Xiao Chen, Shuo-Jin Yang, Shi-Min Hu · 0 citations

The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogene...

Zhen-Tao Tan, Jing-Yi Shen, Yan-Bo Li et al. · 0 citations
#machine learning Preprint Sep 2026

Learning What to Remember: Long-horizon Counterfactual Memory Optimization

Persistent textual memory allows language models to carry information across long interactions, but learning what to remember is fundamentally a credit-assignment problem. A memory rewrite may only become useful many steps later, while much of the observed utility may be inherited from information already stored before...

Jiaming Tang, Ming-Yan Liu, Armin Sarabi · 0 citations
#machine learning Preprint Sep 2026

Constant-Memory Recall: Learned Associations in a Fixed Matrix State

Fixed-size recurrent memory limits storage growth during inference, but successful recall depends on the task and training. We study a small DeltaNet variant with fixed token-specific key biases, trained to remember 32 new key-value pairings per sequence. With 32 KiB of recurrent matrix state, it achieves 99.95% mean a...

Samuel Larson · 0 citations
#machine learning Preprint Sep 2026

Shared Phase and Retention Control for Efficient Adaptive Spectral Recurrence

As new evidence arrives, a sequence model must update what it remembers and how memory influences predictions. While Transformers incur computation and cache costs scaling with context length, fixed-state recurrent models offer constant-memory inference. However, linear and spectral recurrences traditionally rely on st...

Wen-Tao Wang, Heng-Yu Zhong, Yun-Han Jiang et al. · 0 citations
#machine learning Preprint Sep 2026

SchemaMem: Schema-Indexed Recurrent Memory for Delayed State Retrieval

Attention provides direct access to past representations, but retaining an ever-growing history is costly. Recurrent models bound persistent state, yet must preserve selected information while processing subsequent inputs. We introduce SchemaMem, an attention-based recurrent memory architecture combining chunk-local at...

Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.