Skip to content
Preprint

Practical Online KV Cache Compaction for LLM Agents: An Empirical Study

Aug 2026 · 1 citation · 35 references
Computer Science

TL;DR

This work studies online compaction across token eviction (TE) and attention matching (AM), adapting both to compact agent turns and comparing cheap proxy sources such as boundary, repeat-prefill, and delayed future-generation queries, finding TE is often more robust than AM under imperfect proxies.

Abstract

LLM agents accumulate long trajectories of reasoning steps, tool calls, and environment feedback, making the KV cache a major inference bottleneck. KV cache compaction can reduce this cost, but most prior methods assume a static context where future queries are known or can be approximated offline. Agents instead require online compaction: new information must be compressed before future relevance is known, using proxy queries cheap enough for the inference path. We study online compaction across token eviction (TE) and attention matching (AM), adapting both to compact agent turns and comparing cheap proxy sources such as boundary, repeat-prefill, and delayed future-generation queries. Experiments on BrowseComp-Plus and WideSearch show that immediate compaction often hurts performance, whereas delaying compaction to use the agent's future queries recovers much of the gap. Moreover, TE is often more robust than AM under imperfect proxies. Across models at different scales, TE preserves most of the accuracy while reducing KV cache by 80%, and can improve throughput over the no compaction baseline. These results position proxy-query selection as a core design choice for practical online KV compaction.

View source

Similar papers

#machine learning Preprint Sep 2026

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

BeaconKV is proposed, a training-free KV cache compression method that maintains beacon queries, compact representatives for each global query cluster, to anticipate which KV pairs will be revisited without storing the entire query history.

Janghyeon Kim, Minsoo Kim, Kyuhong Shim et al. · 0 citations
Conference Jul 2026

Pegasus: Accelerating Large Language Model Inference with Stateful Prefix Caching

Modern large language model (LLM) inference suffers from severe Time-To-First-Token (TTFT) bottlenecks. Existing prefix KV caching mechanisms are inherently stateless, forcing a trade-off between cross-chunk attention accuracy and online recomputation overhead. To address this issue, we propose Pegasus, a novel statefu...

Fahao Chen, Peng Li, Dongxiao Yu et al. · 0 citations
Preprint Aug 2026

An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age

This work argues that future inference infrastructure should allow decoupling of compute and KV Cache storage across cloud and datacenters, and proposes a vision for an Internet for the KV Cache, with KV Cache management working as a content-distribution system.

Siddhant Ray, Nick Feamster, Junchen Jiang · 0 citations
Jul 2026

Back from the Future: Key-Value Cache Management by Counter-Causal Surprise

Key-value (KV) cache management through compression and eviction strategies has emerged as an important research direction in recent years. Computational demands of large language models (LLMs) and their multi-modal variants during output generation can be partially alleviated by caching previous key and value calculat...

Stephen Gould, A. van den Hengel · 0 citations
#machine learning Preprint Sep 2026

MetaKV: Adaptive KV Cache Compression for Constrained LLM Inference

Key--value (KV) cache compression is an effective way to reduce the memory overhead of large language model (LLM) inference, particularly for long-context workloads. However, existing compression methods make different trade-offs among accuracy, inference latency, and peak KV cache memory utilization, making a single f...

Michael Wang, Keith Li, Roozbeh Bostandoost · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.