Skip to content

An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents

Aug 2026 · 0 citations · 14 references
Computer Science

TL;DR

This work instrument a production compression gateway between coding agents and frontier LLMs, and decompose the token bill of real sessions into three independent levers: tool-schema filtering, content compression of file reads and tool output, and history summarization, which distill the results into an actionable recipe for where token-saving effort pays off.

Abstract

Context compression is widely proposed as a way to cut the token bill of LLM coding agents, and public benchmarks report that aggressive compression preserves task-solving quality. These two facts do not imply the third one commonly assumed: that compressing file reads saves money in a real multi-turn agent. We instrument a production compression gateway (Paritok) between coding agents (Claude Code, Codex) and frontier LLMs (Claude Sonnet, GPT-5), and decompose the token bill of real sessions into three independent levers: tool-schema filtering, content compression of file reads and tool output, and history summarization. Measured in isolation under controlled A/B runs, the three save at fundamentally different rates. Tool-schema filtering removes a fixed block every turn, roughly 21K-57K tokens on a typical turn; it is linear in the turn count N and the only unambiguously and reproducibly positive lever. Content compression saves only about 2% of the cache-priced prefix per turn, but compressed reads accumulate in history and are re-sent on every later turn, so its cumulative saving grows quadratically, about 3350*N^2 tokens (measured), overtaking the fixed tool-filter saving within roughly 6 turns until the context window caps it. A non-destructive gateway lets the agent pull original bytes back on demand; each recall re-sends exactly the one segment just compressed away, so its cost is fixed and bounded rather than a multiplicative blowup, and heavy recall spends the accumulated saving back one segment at a time. Finally, a strong single-shot compression benchmark - 86.5% of SWE-bench quality retained at a 25.7% compression rate, achieved by the model this gateway deploys (Paritok-4B, reported separately) - is orthogonal to multi-turn agent cost and must not be cited as a cost-saving argument. We distill the results into an actionable recipe for where token-saving effort pays off.

View source

Similar papers

Preprint Aug 2026

Paritok-4B: Intent-Conditioned Context Compression for Coding Agents

This work presents Paritok-4B, a 4B LoRA compressor for coding-agent trajectories built on two commitments, and distil a gpt-4.1-mini teacher over 67,074 real OpenHands trajectories into 40,606 validated examples and fine-tune Qwen3-4B.

Jiayu Shi, Lu Chen · 0 citations
#natural language process... Preprint Sep 2026

Protocol Compression Changes Which Party Pays: Bilateral Cost in Cross-Organization LLM Agent Communication

Agents that talk across organizations exchange long messages billed by the token. A shorter notation therefore looks like a saving that costs nothing but an agreement to use it. Recent work reports the saving is conditional. Compressed notation can instead raise total tokens by 8% to 11% over a JSON baseline, when pars...

Janghoon Lee · 0 citations
Preprint Aug 2026

CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents

This study introduces CodeGrep, a 14B retrieval agent trained end-to-end with GRPO to issue multi-turn parallel grep, glob, and read tool calls and return candidate files to a frozen downstream coding agent, and applies the efficiency signal at the advantage layer rather than the reward layer to reduce KL drift and tra...

Wu-Ya Chen, Yihao Yang, Yang Cao et al. · 2 citations
Preprint Aug 2026

The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks

This work reconstructs a coupled-fact graph as reconstructing a coupled-fact graph: at each edit, a required fact comes from recent context or parametric memory, and the facts covered by neither form coherence debt.

Bardia Mohammadi, L. Klein, Aman Chadha et al. · 3 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.