Skip to content

KVpop - Key-Value Cache Compression with Predictive Online Pruning

Jul 2026 · arXiv.org · Vol abs/2607.05061 · 1 citation · 42 references
Computer Science

TL;DR

KVpop is introduced, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision, and introduces a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context.

Abstract

Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly with context length. Existing KV eviction methods often rely on static heuristics or proxy scores, which poorly track future token utility and cause brittle eviction as relevance shifts. To address this, we introduce KVpop, which learns a fixed-budget KV eviction policy by directly supervising the keep-or-drop decision. The scorer is trained against a novel future-attention target, computed efficiently without materializing dense attention maps. We further introduce a delayed memory-based scorer that, uniquely among learned eviction methods, defers scoring for a fixed number of steps to exploit near-future context. On AIME and HMMT mathematical reasoning, KVpop retains 98% of full-attention performance on Qwen3-4B at 75% KV cache compression and 97% at 88% compression, consistently outperforming established eviction baselines. Qwen3-8B shows even stronger results, reaching near-full teacher performance. These results show that supervising eviction with future-attention signals cuts memory costs while maintaining quality.

View source

Similar papers

Preprint Aug 2026

DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

This work proposes DistillCache, a reinforcement learning framework that formulates KV-cache eviction as a sequential decision problem and demonstrates the effectiveness of learned, distribution-aware policies for memory-efficient long-context LLM inference.

Asaad Althoubi · 0 citations
#artificial intelligence Preprint Sep 2026

What Matters for Aggressive Decoding-Time KV Eviction? Temporal Aggregation and Ranking Preservation

Across six open-weight backbones and the LongBench, LongBench-v2, and RULER benchmarks, the results identify temporal aggregation and ranking preservation as distinct, consequential design factors; they do not imply that scoring quality is irrelevant in general.

Bo Zeng, Yu Zhao, Ye-Feng Liu et al. · 0 citations
2025

Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM Inference

A novel method, namely AnDPro, is proposed, which introduces a projection-based scoring function to more accurately measure token importance and guide more accurate token selection in key-Value cache eviction.

Zijie Geng, Jie Wang, Ziqi Liu et al. · 6 citations
Preprint Aug 2026

Sigmoid Attention as a Better Substrate for Learned KV Cache Eviction

Learned KV-cache eviction often faces a soft-to-hard mismatch: during training, differentiable gates typically attenuate token contributions, whereas inference saves memory only when KV entries are physically removed. We ask whether the attention substrate affects this soft-to-hard transition. Using GPT-2-scale Transfo...

I-Hung Li · 0 citations
Preprint Aug 2026

PuzzleKV: Page-Wise Low-Rank Decomposition for KV Cache Compression

Long-context inference in large language models (LLMs) is increasingly limited by the memory required for the key-value (KV) cache. KV cache compression addresses this problem by reducing the storage cost of previous tokens. Among existing approaches, low-rank compression is particularly attractive because it represent...

Zi-Zhong Wang, Jie-Ying Wang, Zhao Zhang et al. · 0 citations
Preprint Aug 2026

QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding

QEvict is proposed, a three-tier KV-cache management scheme that replaces binary retain-or-delete eviction with recoverable eviction, and maintains high-confidence windows in full precision, stores intermediate windows in a quantized recoverable tier, and deletes only the lowest-confidence windows.

Ayushman Garg, Akshita Gupta, Shaswata Bhattacharya et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.