This paper recast context compression as a causal decision preservation problem over discrete interaction units and introduces FOCUS, a training-free context compression framework that operates entirely at test time that requires no offline data collection or fine-tuning.
Abstract
LLM agents accumulate interaction histories that grow linearly with task length, causing quadratic inference cost scaling and performance degradation from attention dilution. Existing context-compression methods learn what to discard offline: by contrastively optimizing guidelines, distilling compressors, or training compression policies. This incurs a substantial cost. Further, the compression policy is learned a priori and is not dynamically conditioned on the evolving test-time trajectories. In this paper we ask a complementary question: Which past interactions causally shape the agent's future decisions? We recast context compression as a causal decision preservation problem over discrete interaction units and introduce FOCUS, a training-free context compression framework that operates entirely at test time. Our method requires no offline data collection or fine-tuning, and is architecture-agnostic, attaching to any closed-API frontier model as a modular compression layer. We evaluate FOCUS on diverse agentic benchmarks including API and tool-calling, QA, web domain and multi-turn dialogue. Our method establishes new state of the art performance, cutting peak context by up to 48% and dependency by 73% while improving task success by up to 8.9 percentage points over uncompressed execution.
This work introduces two exact, gradient-equivalent corrections: LogitTree, a segmented K-forward traversal, and a packed 4D attention mask, and proposes SDCC (Self-Distillation for Conditioning Consistency), a single-backward-pass variational relaxation.
J. Zinco, Xun-Jie Zhu, Shen Huang et al.· 0 citations
This work formalizes CIGD via the drift amplification ratio (DAR), evaluates six compression strategies on a controlled injection benchmark ( per cell), and proposes Goal-Anchored Compression (GAC): a pinned goal anchor, a negation ledger, status-aware retention scoring, and drift re-anchoring.
Zhan Zhang, Wen-Zhi Zhang· Asia Pacific Economic and Ma...· 0 citations
VERA (Visual Evidence-Retaining strategy for long-horizon Agents), a training-free context manager built on deterministic rendering with no exposed memory operations, supporting a modality-preserving view of long-horizon context management.
Jiang-Feng Su, Cong Pang, Jiawei Hong et al.· 1 citation
This work proposes State Conditioned Compression (StateComp), a framework that determines when historical interactions can be safely compressed according to the current agent state, and demonstrates that it reduces total agent and summarization tokens while maintaining task performance, and achieves a 12.67-fold speedu...
Ming-Xuan Wang, Hong-Yue Chen, Ying-Long Guo et al.· 0 citations
A unified taxonomy of agent context compression along three dimensions is introduced: compression target (what is compressed), compression mechanism (how it is transformed and retained), and control policy (who decides when compression is triggered).
It is found that fewer tokens need not mean faster or cheaper execution: on Terminal-Bench with Qwen, policies using roughly one-third as many tokens can take 20-80% longer than the uncompressed agent.
Ritul Satish, Prasoon Sinha, Akiho Kawada et al.· 0 citations