Context Inflation: The Hidden Cost of AI Coding Agents in Real-World Repositories
Abstract
A growing narrative holds that AI coding agents can now do the work of junior software developers, weakening the case for hiring and training them. The evidence offered is almost entirely benchmark performance: leaderboards like SWE-bench report not only how often an agent resolves a task but the dollar cost of each resolution. We argue that this figure does not answer the economic question it appears to. A benchmark task is solved against a frozen snapshot of a repository; real agentic development is continuous work on a codebase that grows, performed by an agent that is stateless, namely it retains nothing across tasks and must re-read the surrounding code on each one. Because the agent carries no memory as the codebase grows, this recurring cost rises with repository maturity — a dynamic we call context inflation — whereas for a junior developer the work of learning a codebase compounds into expertise over a career rather than recurring task by task. This work-in-progress paper introduces a token-level cost model grounded in the documented mechanics of a production agent and applies it to two Python repositories underlying SWE-bench across their full git histories, tracing cost trajectories over a repository’s lifetime as a more ecologically valid basis for comparison than per-task snapshots. Framed this way, the question shifts from whether agents can replace junior developers to who pays for context once they do, and suggests recasting the comparison as one of augmentation rather than replacement.