Skip to content

CliffCompaction: Cost-Efficient Compaction for Long-Horizon Coding Agents

Sep 2026 · 0 citations · 38 references
Computer Science

TL;DR

CliffCompaction is developed, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench and achieving new levels of efficiency for test-time scaling and state-of-the-art results on KernelBench.

Abstract

Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improving performance on Terminal-Bench and achieving new levels of efficiency for test-time scaling and state-of-the-art results on KernelBench. The per-rollout savings of CliffCompaction make the performance--cost trade-off of test-time scaling more efficient, adding over 10 percentage points on Terminal-Bench for less than the cost of two full-context runs. Under parallel test-time scaling, CliffCompaction lets Kimi K2.6 match Opus 4.7, and exceed Opus 4.6 and GPT-5.3 Codex at lower cost. The key to CliffCompaction's effectiveness is that it keeps compacted information faithful by only truncating or dropping content, never rephrasing or rewriting it. We never compact a compaction---each pass operates only on original content, and prior compacted output is discarded, preventing context drift from accumulating. These properties sustain continual learning over sessions exceeding a million tokens: on KernelBench, CliffCompaction reaches CUDA kernel speedups of $2.23\times$ after 200 steps and $3.58\times$ after 400 steps, surpassing specialized search algorithms and trained agents despite being a general-purpose compaction technique. We open-source a scaffold-agnostic API-proxy implementation of CliffCompaction usable with Claude Code, Codex and other harnesses.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

TicTacBench: Benchmarking Timing Closure Capabilities of Coding Agents

TicTacBench is proposed, a benchmark specifically designed to evaluate coding agents' capabilities for RTL-level timing closure under post-place-and-route (post-PnR) evaluation, and TicTacSkill is proposed, a new method that guides agents to follow standard timing-closure procedures and improves the Timing Closure Rate...

Bo-Wei Wang, Zhigang Fang, Zhijie Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Occamy-1.0: Open Pareto-frontier 35B Intelligence for Co-work

Occamy-1.0, a cost-efficient co-work model obtained by further training the post-trained Qwen3.6-35B-A3B checkpoint, is presented and placed at the low-cost knee of the observed cost--performance Pareto frontier.

Wen-Hui Chen, Shi-Wen Cheng, Hao Dong et al. · 3 citations
Preprint Aug 2026

CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing

This paper formulates test-time reasoning as a compute-allocation problem in which a system must decide whether the next unit of compute should be spent on generation, verification, or stopping, and introduces CoBa, a compute-balanced routing policy that first obtains a small set of candidates, applies cheap verificati...

Yan Zhou, Yue Ouyang, Kaiyang Zheng et al. · 0 citations
#natural language process... Preprint Aug 2026

An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents

This work instrument a production compression gateway between coding agents and frontier LLMs, and decompose the token bill of real sessions into three independent levers: tool-schema filtering, content compression of file reads and tool output, and history summarization, which distill the results into an actionable re...

Lu Chen, Jiayu Shi · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.