Jul 2026
Reward-Free Evolving Agents via Pairwise Validator
A pairwise validator is proposed: a frozen LLM that, given the parent and child candidate, returns a binary verdict on which is better, which mitigates the need for strict scale calibration.
Minghao Liu, Yu Wang, Jiayun Wang et al.
· arXiv.org · 0 citations