Skip to content

Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation

Jul 2026 · arXiv.org · Vol abs/2607.08938 · 5 citations · ⚡ 2 influential · 59 references
Computer Science

TL;DR

This work creates a framework that maps agent failure modes to harness adaptation strategies, and builds a harness optimizer that automatically discovers effective adaptations from failure trajectories, suggesting that harness adaptation can expand the practical deployment range of SLM agents in routine business tasks.

Abstract

Frontier LLM agents are automating many business tasks, but their high inference cost makes large-scale deployment unsustainable. Small language models (SLMs) offer a cheaper alternative, yet they typically fall short when swapped into a harness designed for a frontier LLM. We show that for many routine business tasks, SLM agents can match LLM performance at 90% lower cost, when paired with an adapted harness that can be automatically discovered by a meta agent. The key insight is that much of the task difficulty is shared across instances and can be lifted from the model into the harness via tailored instructions, tools, and orchestration loops. To study this systematically, we create a framework that maps agent failure modes to harness adaptation strategies, and build a harness optimizer that automatically discovers effective adaptations from failure trajectories. Across seven business-oriented agentic tasks and three SLM families, we found optimized harnesses significantly improve performance on 16 of 21 task-SLM pairs, with seven pairs closing the SLM-LLM performance gap and the best SLM agent recovering 89.7% of LLM performance at 4% of the cost. Our analysis further shows that adaptation works best for tasks with more repetitive workflows and for SLMs with sufficient base capabilities. Together, these results suggest that harness adaptation can expand the practical deployment range of SLM agents in routine business tasks.

View source

Similar papers

Preprint Aug 2026

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

The proposed AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches, suggests that automatic harness optimization is a promising path toward more performant and reliable agent...

Sungho Park, Wonjoong Kim, Rongyuan Tan et al. · 14 citations
#artificial intelligence Preprint Sep 2026

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

This work proposes Ecdysis, which aggregates failure evidence across task instances before promoting recurring failure patterns into persistent harness evolution, biasing evolution toward repairs that are more likely to generalize beyond individual model behaviors.

Rui-Qing Yue, Yu Cui, Zhuo-Yu Sun et al. · 1 citation · ⚡1
#natural language process... Preprint Sep 2026

HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?

It is found that generated harnesses remain substantially behind mature human-engineered references on code and on search and research, while matching or exceeding the selected references on writing and machine-learning experimentation, with large variation in execution cost.

Yu-Hao Wu, Jing-Yuan Zhang, Jia-Jun Shi et al. · 8 citations
Preprint Aug 2026

Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification

HarnessLens is introduced, a budget-aware framework for automated harness evolution that jointly explores the task space and user-configurable components, derives candidate modifications from execution trajectories, and selectively verifies each candidate on behavior-relevant tasks using an attributable-evidence gate.

Jing-Heng Xu, Yi-Kai Zhang, Aiden Chen et al. · 7 citations · ⚡1
Preprint Aug 2026

Evo-Bench: Can Language Models Improve Agent Harness?

Evo-Bench is the first benchmark designed to evaluate models'intrinsic harness-evolving capabilities across Search, Office, and General agent domains, and exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consis...

Lisheng Huang, Chen Yang, Hao Zhou et al. · 6 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.