Skip to content

Reward-Free Evolving Agents via Pairwise Validator

Jul 2026 · arXiv.org · Vol abs/2607.14408 · 0 citations · 34 references
Computer Science

TL;DR

A pairwise validator is proposed: a frozen LLM that, given the parent and child candidate, returns a binary verdict on which is better, which mitigates the need for strict scale calibration.

Abstract

A self-evolving agentic loop repeatedly proposes a tweaked version of an agent (its prompt template or program) and accepts or rejects the change based on a per-iteration quality signal. Designing that signal is often the costly part of the project: a reliable scalar reward requires domain expertise and labeled examples that are themselves as expensive to assemble as the agent's underlying task. We propose replacing the scalar at the accept/reject gate with a pairwise validator: a frozen LLM that, given the parent and child candidate, returns a binary verdict on which is better. Pairwise judgment is generally easier and more stable than absolute scoring, due to its contrastive nature, which mitigates the need for strict scale calibration. The validator also requires no training of its own. We integrate the validator into three published self-evolving engines (GEPA, ADRS, ShinkaEvolve) and report two flavors: Adaptive Focus, which retains the engine's existing val-set parent selection, and Soft Elo, which lets the validator's verdicts drive parent selection so that val-set rewards drop as well. Across multiple agents and two artifact substrates (prompt and code), our method matches or exceeds the full-reward baseline on the majority of settings we evaluate, and the pattern survives a cross-family validator swap. The pairwise gate is thus a drop-in replacement for per-step reward design at competitive task accuracy without the labeling cost.

View source

Similar papers

Preprint Jul 2026

DarwinX: Evolving Agent Harnesses Through Natural Selection

DarwinX is introduced, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence sh...

Yifang Zhang, Yutong Dai, Juntao Tan et al. · 2 citations · ⚡1
#artificial intelligence Preprint Aug 2026

Self-Evolving Skills via Surrogate-Guided Solve-and-Reproduce

Agent skills are portable packages of instructions and resources an agent consults at deployment. Self-evolving them fails in two ways today. First, skills evolved from scratch underperform human-curated ones and, on a weak model, using no skill at all. Second, an evolution-time pass records one lucky trajectory that a...

Jia-Le Liu, Pin-Ze Ren, Yu Xia et al. · 0 citations
Preprint Aug 2026

Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?

This work introduces Open-Ended Optimization (OEO), which keeps the objective, permitted interactions, resource budget, data boundary, and evaluation fixed while allowing the optimizer to compose the improvement process online.

Xue Hui, Fan Yang · 1 citation
Jul 2026

SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents

The proposed SeekJudge framework, in which four role-specialized agents, a Condense, a Ground, a Seek and an Analyze agent, reach a verdict through a Seek--Analyze loop over the trajectory, is the first practical model-based reward to match or surpass native rule-based supervision in online RL.

Yang Wan, Zhenhao Zhang, Jie-Rui Wang et al. · 0 citations
Preprint Aug 2026

Envs-FORGE: Frontier-Optimized Reward-Grounded Environment Synthesis for Agent RL

Reinforcement learning (RL) for terminal agents needs executable training environments with reliable rewards and useful difficulty. Fixed recipes such as few-shot, Self-Instruct, and Evol-Instruct apply the same prompting policy to every seed, even when the current policy would benefit from a harder, easier, or simply...

Xiao-Jun Wu, Ce-Hao Yang, Honghao Liu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Belief-Calibrated Optimization: An Explicit World Model for Agentic Optimization

The performance of an LLM agent depends on the scaffold around a frozen model. A common way to improve that scaffold is to use a coding agent as an optimizer: it reads current scores and traces and iteratively edits the source, producing a new candidate each round. Each edit is chosen according to a belief about how th...

Yu-Han Chen, Zhi-Hua Tian, Mahavir Dabas et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.