Skip to content
Preprint

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

Aug 2026 · 2 citations · 38 references
Computer Science

Abstract

We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. Existing approaches, however, evaluate a fixed validation set in full at every iteration, incurring substantial evaluation costs even on tasks that become less discriminative as the harness evolves. We propose $\textbf{Task-CoEvolve}$, which co-evolves the validation tasks with the harness by addressing two challenges: selecting informative tasks and estimating full-set performance from partial evaluations. Task-CoEvolve builds on the observation that tasks on which candidate harnesses disagree are more informative for distinguishing among them than tasks that are consistently solved or failed. It uses variance-weighted sampling based on past outcomes to focus evaluation on tasks near the capability frontier, with the sampling distribution adapting as the harness evolves. It then estimates full-set scores from the sampled tasks by accounting for their sampling probabilities, enabling consistent comparisons across iterations despite evaluating different subsets. Experiments on online text classification and Terminal-Bench 2.1 show that Task-CoEvolve consistently outperforms subset-based baselines and matches the final performance of full-set search while reducing the number of evaluations during optimization by 80%. Code will be released at https://github.com/Agent4Science-UTokyo/Task-CoEvolve.

View source

Similar papers

Preprint Aug 2026

Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets

Janus is introduced, a framework that uses LLMs to co-evolve target programs and executable proxy evaluators to address label scarcity and extend evaluator-guided LLM discovery from tasks with cheap, scalable feedback to scientific domains where trustworthy evaluation is scarce and expensive.

Xi-Meng Liu, Qianlong Wang, Ying-Ming Mao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Direct Self-Evolving Optimization: Evolving LLMs without Challenger Training

Self-evolving language models improve by generating tasks and learning from their own feedback, but adapting the task generator often requires a separate challenger-training loop. Can we generate tasks adapted to the current solver without explicitly training a challenger? We introduce \textbf{D}irect Self-\textbf{E}vo...

Yu-Yang Deng, Yu Wang, Jia-Yun Wang · 0 citations
#artificial intelligence Preprint Sep 2026

VACE: Validation-Gated Alternating Co-Evolution of Agent Models and Harnesses

Language model agents can be improved by updating their model weights or refining the harness that guides task execution. These components are coupled: weight updates change how the model uses the harness, while harness updates change the trajectories used for training. We propose VACE, Validation-Gated Alternating CoE...

Jie-Xing Qi, Yu He, Jun Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance Ranking

As the complexity of High Performance Computing (HPC) ecosys- tems continually increases, achieving optimal performance becomes a challenge. Traditional performance autotuning techniques pro- vide promising means to navigate this complexity, these techniques remain computationally intensive and require many evaluations...

Md Arafat Hossain, Thomas Randall, Akashnil Dutta et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MoMHa: Multi-Objective Optimization of LLM Harnesses over Accuracy, Safety, and Tokens

Most work on improving large language models treats accuracy as the sole objective. We argue that the harness, the Python code surrounding the model that constructs prompts, routes calls, and parses outputs, is a first-class design surface whose quality is inherently multi-objective: an accurate harness that refuses no...

Subhojyoti Mukherjee, Mahmud Tanjim · 0 citations
#artificial intelligence Preprint Aug 2026

Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck

The first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing is conducted - grounded in a unified framework that decomposes the effectiveness of each method's token budget into exploration and exploitation.

Davide Romano, Kanak Raj, Jerrod Parker et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.