Skip to content

When Does Restricting a Coding Agent to execute_code Help? A Regime ⨉ Agent-Design Ablation

Jul 2026 · arXiv.org · Vol abs/2607.10569 · 0 citations · 46 references
Computer Science

TL;DR

Two implications: the cheapest tool surface is jointly determined by task regime and agent design rather than by either axis alone, and the headline cost signal lives in cache-adjusted cost -- not pass rate, which is invariant across surfaces at the model sizes the authors evaluate.

Abstract

Modern coding agents expose multiple tool surfaces -- IDE primitives, bash, and Model Context Protocol (MCP) code-execution -- and the field has shipped three contradictory claims about which one matters. We run the missing crossed comparison: an integrity-clean three-arm ablation (baseline / bash_only / code_only) on synthetic computation tasks and SWE-bench Mini modification tasks, holding model, harness, and prompts fixed, with two agents (Claude Code, OpenAI Codex CLI) so the comparison spans both regime and agent-design axes. Across the four resulting (regime, agent) cells, restricting the agent to a single execute_code MCP tool is cheaper than -- or statistically tied with -- its cheapest tool-rich rival in three cells (significantly on Artifact/Claude and SWE-bench/Codex; directionally on Artifact/Codex), with pass rates statistically tied within each cell. The lone exception is SWE-bench/Claude, where code_only is directionally costlier (+14.4%, not significant); a conditional-cost analysis localizes that gap to failure-cost on doomed-run trajectories, not a per-edit tax on successful runs. Two implications: the cheapest tool surface is jointly determined by task regime and agent design rather than by either axis alone, and the headline cost signal lives in cache-adjusted cost -- not pass rate, which is invariant across surfaces at the model sizes we evaluate. The benchmark harness, task suite, and analysis code are available at https://github.com/hyang0129/onlycodes.

View source

Similar papers

Preprint Aug 2026

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

This work introduces Harness-IF, which scores operational rules one at a time from execution evidence: 60 realistic multi-turn coding items drawn from a 642-rule library, 256 rules receiving verdicts, placed on the five configurable surfaces a deployed agent reads.

Zining Huang, Haoran Que, Hongxia Zeng et al. · 1 citation
#artificial intelligence Preprint Aug 2026

A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations

This work introduces a random variant sampler that applies common semantics-preserving transformations (SPTs) - spanning control-flow rewrites, dead-code injection, and identifier renaming - to produce perturbed variants, demonstrating that even top frontier models are susceptible to semantics-preserving perturbations.

Hasan Mahmud, Shreya Gupta, Isha Chaudhary et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite

An agentic coding system couples a language model to a harness: the tools, prompts and control flow that turn a chat model into an autonomous software engineer. Vendors ship harnesses tuned to their own models, and practitioners assume the vendor-native pairing solves more tasks. We measure that assumption with paired...

Mohsen Arjmandi · 0 citations
Preprint Aug 2026

Same Model, Different Harness: Different Coding-Agent Results

A coding agent combines a model with a harness, which decides what the model sees, which tools it can use, and how the work continues, and two configurations of the same harness are compared on three coding benchmarks.

S. Lewis · 2 citations
#artificial intelligence Preprint Sep 2026

An Empirical Study of Harness Design for Coding Agents

Coding harnesses shape how autonomous coding agents translate model capabilities into long-horizon software-engineering performance, yet existing work typically evaluates harnesses as monolithic systems, leaving the effectiveness of individual components unclear. To enable component-level comparisons, we study this que...

Run-Ze Fan, Zi-Hao Zhang, Si-Min Ma et al. · 1 citation · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.