Beyond Token Savings: A Systematic Study of Context Compression in LLM Agents
It is found that fewer tokens need not mean faster or cheaper execution: on Terminal-Bench with Qwen, policies using roughly one-third as many tokens can take 20-80% longer than the uncompressed agent.
Ritul Satish, Prasoon Sinha, Akiho Kawada et al.
· 0 citations