PREreview of "Control the Harness, Control the Cost: Routing and Governing AI Coding Agents in the Enterprise"
Abstract
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/22974223. Inside a coding-agent harness one user turn is many API requests on a prompt cache, and the authors route only where no live conversation needs that cache rebuilt. In an emulated enterprise of 10,000 seats this recovers 14 to 21% of model spend. What I liked - the payback rule is an equation (Equation 1), Table 4 changes one assumption at a time, and Section 8 is upfront about limits (synthetic scenarios, list prices only). I read it as a practitioner who runs coding agents every day. In my company software development is done by AI agents and I measure what they cost me, so maybe my practical side is useful. My four comments: where the spend sits, clearing context, the cache-read share, confidence and checks. 1. Regarding where the spend sits - Figure 3 shows that for both corpora the top 1% of sessions carry 31 to 53% of spend, and my logs point the same way. I measured 722 agent sessions (150,902 model calls, 34.6 billion tokens) and about 3% of sessions gave 80% of the spend, sessions longer than 200 calls (8% of sessions) gave more than 90% of the money (DOI 10.5281/zenodo.22759216, technical report, not peer-reviewed). These are sessions I ran myself, cost modeled from list prices, not bills - the same kind of repricing you use. My question is about Section 8 - the 51 to 80% share of long sessions that are cheaper on Fable 5.1 rests on only 102 TraceLab and 30 SWE-chat sessions that cost $100 or more. An interval around that share would show how much the headline leans on this small tail. 2. About clearing context - Section 6.1 says clearing old reasoning and tool results rewrites the cache, so it rarely pays on a warm cache and is nearly free on an expired one. I tested how often to clear context between tasks - 36 runs, six policies, six repeats. The curve was U-shaped, the minimum at three tasks per session, three to six a plateau, and clearing after every task cost about a third more (+33.5%, p = 0.002, 0.011 after Holm correction, DOI 10.5281/zenodo.22759217). Two caveats - the comparison point was picked after the runs, and the clear-every-task arm wrote more tests, so part of the gap may be extra work. That second one is a quality effect of the kind Section 8 leaves outside the emulation. When the pilot group from Section 6.1 tests a context standard, I'd count the work too (tests written, corrections) next to the spend. 3. On the cache-read share - your 98.1% in Section 8 sits above TraceLab (95.2%) and SWE-chat (97.6%), and you say this favours the crossover. In my own 168-session run (DOI 10.5281/zenodo.22712985) 94.4% of paid tokens were re-reads of context already sent. It's a share of tokens, not of the bill, and not the same metric as your share of input read from cache, so only a rough comparison - but it sits below both corpora. The saving moves a lot with this share and with one price line, the Fable 5.1 cache-read discount (Table 4: 5.2% under a uniform 0.1x cache-read price). A plot of saving against the cache-read share, plus a line on the price dependence in the abstract, would let a reader place their own workload. 4. About confidence and checks - I agree with not trusting stated confidence (Section 4.6 routes on calibrated probabilities, not the model's own confidence). In my work the model sounds the same when it's right and when it's wrong, and every task needs a check the agent can run itself, otherwise "looks done" is the only signal the agent has. I use a second, fresh agent whose task is to disprove the first; independent agents rarely make the same mistake, but one round lowers the risk, it doesn't remove it, and the important things a person checks. As I read the emulation, a correction is charged only when a user notices the problem, and a wrong answer that sounds sure may pass without one. A verification step per task (a test or a second-agent check) would make the cost per verified task from Section 6.1 measurable in the case study. Competing interests Yes: the text cites the author's own technical reports (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22759217, 10.5281/zenodo.22712985); no connection to the article's authors. Use of Artificial Intelligence (AI) The author declares that they did not use generative AI to come up with new ideas for their review.