Real-world users often exhibit highly heterogeneous preferences over multiple objectives for LLM responses. A lightweight aligner can tailor these responses to individual preferences, but scarce user-specific feedback makes personalized training difficult. Learning shared initializations across users can support few-sh...
An extensive evaluation of automatic harness evolution for LLM agents is conducted, comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluating evolved harnesses on held-out tasks to assess whether the discovered improvements gen...
Yike Wang, Huaisheng Zhu, Zhengyu Hu et al.· arXiv.org· 18 citations
Co-Harness is introduced, a framework that jointly optimizes the agent harness and model parameters during post-training and suggests that joint harness and model optimization is an effective way to improve agents beyond fixed-harness post-training.
Zhengyu Chen, Teng Xiao, Huaisheng Zhu et al.· arXiv.org· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.