Skip to content

Automated Discovery Has No Universally Superior Harness

Jul 2026 · arXiv.org · Vol abs/2607.18235 · 2 citations · 50 references
Computer Science

TL;DR

This work systematically decomposes OpenEvolve-style evolutionary search and the TTT-Discover search harness into its constituent components and systematically evaluates 30 budget-matched harnesses across 12 model-problem pairs using more than 3.1 million LLM rollouts and repeated-trial statistical analysis, showing that discovery harnesses have a generalization problem.

Abstract

Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these are composite systems combining several design choices about archives, parent selection, exploration, and budget allocation into a single recipe. Because discovery runs are expensive and inherently stochastic, existing harnesses are often compared using too few independent trials to distinguish key methodological improvements from run-to-run variance. We systematically decompose OpenEvolve-style evolutionary search and the TTT-Discover search harness into its constituent components and systematically evaluate 30 budget-matched harnesses across 12 model-problem pairs using more than 3.1 million LLM rollouts and repeated-trial statistical analysis. Our results show that discovery harnesses have a generalization problem: No fixed harness is reliably superior across the evaluated model-problem pairs, and variants of OpenEvolve generally underperform simpler alternatives. Thus, harness choice is better viewed as a hyperparameter rather than as a universal recipe, and should be tailored to the specific problem and underlying model. We also find that early discovery progress predicts final performance, and use this property to present a budget-matched adaptive-allocation experiment that starts multiple harnesses, prunes weak partial runs, and reallocates compute to stronger survivors, outperforming both commitment to a randomly sampled fixed harness and a non-adaptive harness ensemble. Together, these results motivate shifting from fixed harness selection to online adaptation guided by early performance. We release all run pools including baseline null distributions for every model-problem pair as reusable statistical infrastructure against for future harness proposals.

View source

Similar papers

Preprint Aug 2026

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

The proposed AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches, suggests that automatic harness optimization is a promising path toward more performant and reliable agent...

Sungho Park, Wonjoong Kim, Rongyuan Tan et al. · 14 citations
Preprint Aug 2026

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

This work evaluates 5 frontier LLMs as optimizers both under a shared coding harness and under their native harnesses across 4 downstream tasks, and establishes harness optimization as a measurable and discriminative capability with large space for improvement.

Varun Ursekar, Apaar Shanker, Yash Maurya et al. · 5 citations
#artificial intelligence Preprint Sep 2026

Direct Optimization of Generators for Search in Automated Theorem Proving

This work extends Compute-Aligned Training to this setting through an abstraction of policy-guided search, deriving tractable, trace-supported losses and introduces a search-agnostic uniform-allocation (UA) loss that accounts for the budget without specifying the specific search.

Adam Ousherovitch, A. Tewari · 0 citations
Book Open access Sep 2026

Which LLM to Fine-Tune? Agent-Driven Model Selection at Scale

It is shown that model selection is a recommendation problem, and AgentRec, a multi-stage retrieval-and-ranking framework that progressively narrows hundreds of candidate models using increasingly expensive but more faithful evaluation signals, is introduced.

Chen Luo, Yu-Lin Liu, Yi Liu et al. · 0 citations
Preprint Aug 2026

GenomeHarness: Harnessing Al Agents for Reliable Adaptation of Genome Language Models

GenomeHarness, an agentic harness for adapting genome language models through controlled search over fine-tuning recipes, improves mean test MCC in 47 settings, and shows gains on Genomic Benchmarks and on tasks where the root recipe is unstable or poorly matched.

Weicai Long, Yusen Hou, Houcheng Su et al. · 0 citations
Preprint Aug 2026

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. Existing approaches, however, ev...

Atsuyuki Miyai, Kiyoharu Aizawa, Toshihiko Yamasaki · 5 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.