Skip to content
Preprint

Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding

Aug 2026 · 2 citations · 43 references
Computer Science

TL;DR

The resulting divergence catastrophic remembering is named: the inverse of catastrophic forgetting around which continual learning is organized, the inverse of catastrophic forgetting around which continual learning is organized.

Abstract

Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrites the file wholesale. We trace this to imperfect recall: appending an instruction is always cheap, but once an instruction's rationale is gone, deleting it without risking a correctness regression costs O(2^|D|) in a prompt of |D| instructions. We name the resulting divergence catastrophic remembering, the inverse of catastrophic forgetting around which continual learning is organized. First, we characterize this phenomenon across 247,694 instruction lifetimes in 1,867 repositories: agentic prompts grow without bound, more than tripling over their lifetime (+226%), gaining +4.9 net instructions every commit; further, the older an instruction gets, the less likely it is to be deleted (log-hazard -0.032/commit). Then, we show that prompt comments can halt the growth: inverting IFEval yields verifiable worlds whose optimal prompts are known, and there comments encoding latent reasoning remove 99.3% of excess instructions (+211.3% to +1.4%). Finally, applying the same inversion to WildIFEval, we show that prompt comments can improve real-world agentic instruction-following by up to 23.1%. If English is the new code, why don't we have comments yet?

View source

Similar papers

Preprint Aug 2026

Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression

Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being written to memory. We ask what governs whether it does. Matched notes carry the identical claim and identical stance and differ only in where that stance sits; one model...

Alex Kwon · 0 citations
#machine learning Preprint Sep 2026

MemoryWalker: Stop Training Agents on Contexts They Never Saw

This work introduces two exact, gradient-equivalent corrections: LogitTree, a segmented K-forward traversal, and a packed 4D attention mask, and proposes SDCC (Self-Distillation for Conditioning Consistency), a single-backward-pass variational relaxation.

J. Zinco, Xun-Jie Zhu, Shen Huang et al. · 0 citations
Preprint Aug 2026

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

This work introduces Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads.

Zining Huang, Haoran Que, Hongxia Zeng et al. · 1 citation
Preprint Aug 2026

Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents

The results suggest that effective long-horizon agent memory depends less on storing more information than on deciding which information should remain active, and that effective long-horizon agent memory depends less on storing more information than on deciding which information should remain active.

Q. Dao, Purvi Kathalkar, Kenneth Eaton · 1 citation
Jul 2026

Presentation, Not Mechanism: A Render Confound in Deprecation-Aware Memory Evaluation

AI systems increasingly retrieve from records that revise themselves: issue threads, encyclopedic histories, policy logs, and long conversations. The challenge is not only finding relevant evidence, but deciding which claims remain in force, which were superseded, and when to abstain. Structured memories promise to sol...

Zhao-Yang Jiang, Zhizhong Fu, Zicheng Li et al. · 0 citations
Preprint Aug 2026

Causal Episodic Memory for Feedback-Driven Agent Repair

Results clarify when causal cross-query memory improves repair and when broader memory representations remain preferable, and show that negative memory contributes modestly, the value of type conditioning and lexical-dense ranking is dataset dependent, and schema-local experience provides the most consistent benefit.

K. Vo, Tam Minh Chu, A. T. D. Dinh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.