Skip to content
Preprint

The Compaction Cliff in Long-Running AI Agent Memory

Aug 2026 · 0 citations · 41 references
Computer Science

TL;DR

Knowledge Triage, a framework that classifies each line of an agent's knowledge base by type and routes each type through its own retention policy, is addressed, and AgentArtifactCorpus, the classifier, and the reference implementation are released.

Abstract

A safety rule and an episodic log compete for the same tokens in an AI agent's context. When the budget overflows, both are summarized at the same rate; only the rule needs exact wording to remain enforceable. On 20 production agent configurations, Claude Code's /compact prompt on Sonnet 4.6 preserves 53\% of safety rules after one compaction round and 10\% after five. We name this the Compaction Cliff. We address it with Knowledge Triage, a framework that classifies each line of an agent's knowledge base by type and routes each type through its own retention policy. Three deterministic operators implement this triage across the three context-management operations: TypeCompact rewrites items in place under per-type fidelity, TypeDecompose partitions a topic too large to compact safely, replicating in-scope safety rules across partitions, and TypeRetrieve fetches items from external storage with in-scope rules pinned ahead of relevance. On five public corpora, TypeCompact preserves 2--4$\times$ more safety rules than the strongest single-shot LLM compactor at every ratio, with 96\% recall over five rounds. TypeDecompose reaches 0\% locality violations against 93\% under uniform partitioning. TypeRetrieve reaches 100\% recall@50 against 73\% for the best single-shot LLM retriever. On three downstream behavioral benchmarks, we outperform the production Sonnet compactor on medical compliance (paired McNemar $p<10^{-8}$ on preservation, $N = 200$), the full-policy and hierarchical baselines on retail task pass rate ($p<0.01$, $N = 115$), and the hierarchical compaction on the airline domain ($p = 0.024$). We release AgentArtifactCorpus (396{,}934 agent configurations from 54{,}628 public GitHub repositories), the classifier, and the reference implementation.

View source

Similar papers

Preprint Aug 2026

Muscle Memory for Agents: Compile not Merely Retrieve

This paper argues that Muscle Memory - the practice of compiling recurring user intent into purpose-built specialist agents - is a distinct memory paradigm from retrieval, and argues that compilation is a better fit for the workloads where current assistants impose a multi-turn tax on their users.

Pouya Ghiasnezhad Omran, Soujanya Lanka, Qin Zhang et al. · 0 citations
Jul 2026

Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory

This work decomposes each operator's utility into a coverage effect on evidence omitted by retention and a signed replacement effect on raw evidence that already fits, which explains why the preferred action changes with relative budget pressure.

Qingcan Kang, Mingyang Liu, Shixiong Kai et al. · 1 citation
Jul 2026

Addressable Recall Compaction for Long Context-Window Control in AI Agents

Results indicate that explicit, address-based recall can improve information retention and serving efficiency relative to the evaluated context-management baselines under the tested settings.

T. Dang, Yuma Ichikawa, Sakina Fatima et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Plan Pointers and Record-Directive Form in Budgeted Verification of Inherited Agent Memory

A model that inherits one-line memories may pull one archived source record before acting; a directive in the store can steer that pull: a pointer, a criterion or both. Across sixteen registered studies (179,352 attempts) we measured where the request goes under each form; every result is descriptive, with registered i...

Kazuki Nakayashiki · 0 citations
Preprint Aug 2026

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

This work introduces Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads.

Zining Huang, Haoran Que, Hongxia Zeng et al. · 1 citation
#natural language process... Preprint Sep 2026

PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents

PARSER, which decouples reading from reasoning, is introduced, which is robust to perturbations in evidence position, order, and distance, conditions that cause large accuracy swings in sequential methods, while reducing inference latency by up to 11x.

Kun Li, Ze-Xuan Qiu, Tian-Hua Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.