CompKV is introduced, the first compensation-aware sparse attention framework that divides tokens into blocks and explicitly optimizes selection for the downstream compensation mechanism, and shows that the residual left by block-level mean compensation is governed by both block attention mass and within-block logit variation.
Abstract
Despite their strong performance, large language models (LLMs) are bottlenecked by KV cache memory traffic during long-context inference. Sparse attention is widely used to accelerate LLM inference by computing exact attention over a selected subset of tokens. To recover the contribution of tokens excluded from exact attention, recent methods apply coarse-grained compensation to the omitted attention tail. However, existing methods typically select tokens based on attention mass and only then compensate for the unselected tokens. This decoupled design overlooks their interaction: selection should prioritize tokens that would leave the largest compensation error if omitted. To address this limitation, we introduce CompKV, the first compensation-aware sparse attention framework that divides tokens into blocks and explicitly optimizes selection for the downstream compensation mechanism. Our theoretical analysis shows that the residual left by block-level mean compensation is governed by both block attention mass and within-block logit variation. We approximate this residual using compact block-level statistics, yielding a deployable selection criterion. We further develop an efficient asynchronous implementation. Experiments on RULER and LongBench-Pro show that CompKV performs best among the evaluated sparse baselines while delivering up to a $6.85\times$ self-attention speedup over full attention.
QEvict is proposed, a three-tier KV-cache management scheme that replaces binary retain-or-delete eviction with recoverable eviction, and maintains high-confidence windows in full precision, stores intermediate windows in a quantized recoverable tier, and deletes only the lowest-confidence windows.
Ayushman Garg, Akshita Gupta, Shaswata Bhattacharya et al.· 1 citation
A novel method, namely AnDPro, is proposed, which introduces a projection-based scoring function to more accurately measure token importance and guide more accurate token selection in key-Value cache eviction.
Zijie Geng, Jie Wang, Ziqi Liu et al.· Neural Information Processin...· 6 citations
This work proposes DistillCache, a reinforcement learning framework that formulates KV-cache eviction as a sequential decision problem and demonstrates the effectiveness of learned, distribution-aware policies for memory-efficient long-context LLM inference.
Large language model (LLM) inference is often constrained by both computation and memory, especially in offloading-based deployments where model weights are transferred across memory hierarchies during autoregressive decoding. In this setting, reducing the number of executed layers can lower per-token latency while als...
Long-context inference with Large Language Models (LLMs) is bottlenecked by the linearly growing memory of the key-value (KV) cache. Existing compression methods reduce the cache through token eviction or approximation, but degrade sharply at aggressive compression budgets. We propose PatchKV, a training-free framework...
Chanryeol Lee, Chanhyuk Lee, Yeonwoo Choi et al.· 0 citations
Key-value (KV) caching is essential for efficient autoregressive large language model (LLM) inference, but the cache grows linearly with context length, increasing storage and decoding costs. KV cache compression mitigates this cost by retaining only a subset of cached tokens. This challenge is particularly important f...
Bo-Yu Feng, Jia-Hong Liu, Yi-Fan Li et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 29, 2026
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.