Skip to content

Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks, it matches the strongest baseline in task performance while delivering 32-43% higher throughput than that method when deployed with vLLM.

Abstract

Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some estimate of how much it will matter later, and keep the top-scoring ones. We show that the selection signal contributes almost nothing. Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks, it matches the strongest baseline in task performance while delivering 32-43% higher throughput than that method when deployed with vLLM. Controlled experiments explain this by showing that 1) the prompt is the fragile part of the cache, and most of the gap between selectors is just whether their selection signal happened to keep it; 2) the reasoning trace protects itself against eviction with redundancy at two levels, in the text (the model restates what it still needs as it works) and across attention heads (each keeps its own copy of the trace), so once the prompt is safe, a random draw retains enough copies of what the model still needs, and no score is required to pick them. Our code is publicly available at https://github.com/SalesforceAIResearch/Random-Attention.

View source

Similar papers

Preprint Aug 2026

DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

This work proposes DistillCache, a reinforcement learning framework that formulates KV-cache eviction as a sequential decision problem and demonstrates the effectiveness of learned, distribution-aware policies for memory-efficient long-context LLM inference.

Asaad Althoubi · 0 citations
#machine learning Preprint Aug 2026

PAGE: Partition-Aware Gated KV-Cache Eviction

KV-cache eviction can do more than compress. In long-context LLMs, keeping only some cached tokens sometimes matches or exceeds full-cache accuracy, because many redundant prefill tokens otherwise dilute attention away from the tokens that carry the answer. This benefit is not uniform, and evicting the wrong tokens can...

Pankaj Kumar, Subhankar Mishra · 0 citations
Preprint Aug 2026

Runtime Observability for Heterogeneous Attention Memory

A runtime observability contract is given that covers all four memory classes with three operators, instantiate it on six model configurations across five architecture families, and compose the per-stage bounds into an executable request-level risk ledger.

Fanzhe Wei, Li Liu, Ziyang Wang et al. · 2 citations
#machine learning Preprint Sep 2026

When to Evict, Not What to Keep: Draft-Guided Eviction for Training-Free KV-Cache Compression

Training-free KV-cache compression methods such as SnapKV, H2O, and PyramidKV evict tokens at the end of prefill, aiming to preserve the attention mass that future queries are expected to use -optimizing what to keep. We show that this objective fails in two distinct ways. (1) Compensation: restoring the evicted attent...

Haeyong Kang, C. D. Yoo · 0 citations
#artificial intelligence Preprint Sep 2026

CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference

CompKV is introduced, the first compensation-aware sparse attention framework that divides tokens into blocks and explicitly optimizes selection for the downstream compensation mechanism, and shows that the residual left by block-level mean compensation is governed by both block attention mass and within-block logit va...

Zheng-Hong Huang, Rui-Zhe Yao, Dan-Yi Liu et al. · 1 citation · ⚡1
Preprint Aug 2026

KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving

KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history. Under aggressive budgets, this not only lowers accuracy but can also cause runaway degeneration, where the model produces incoherent or repetitive tokens until reaching t...

Minsoo Cheong, Woo-Sang Lim, Vincent-Daniel Yun et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.