Practical Online KV Cache Compaction for LLM Agents: An Empirical Study
This work studies online compaction across token eviction (TE) and attention matching (AM), adapting both to compact agent turns and comparing cheap proxy sources such as boundary, repeat-prefill, and delayed future-generation queries, finding TE is often more robust than AM under imperfect proxies.