While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems remain bottlenecked on pre-defined and hand-crafted model pools. In this work, we investigate the problem of model selection in multi-LLM systems. We propose and systemat...
Zong-Wan Cao, Zi-Yuan Yang, Shang-Bin Feng et al.· 0 citations
A pool of language models can collaborate and improve collectively by learning from one another's responses. These interactions depend on the instructions used during training. Existing methods typically sample instructions uniformly, even though their usefulness may change as the models improve: an instruction on whic...
Christina Hahn, Shang-Bin Feng, Dean Light et al.· 0 citations
Long-horizon coding agents receive verifiable rewards only after completing expensive sequences of tool calls. This increases inference cost, amplifies early wrong hypotheses, and can lead to sparse terminal reward and unstable training. We introduce Contextual Early Reward (CER), which predicts terminal reward through...
Ji-Han Yao, Si-Han Zeng, Shang-Bin Feng et al.· 0 citations
Group-relative reinforcement learning (RL) relies on reward variation among sampled responses to estimate informative relative advantages. As language models become increasingly capable, existing training data can become reward-saturated: all sampled responses to the same problem might receive equally high rewards, whe...
Zi-Yuan Yang, Yike Wang, Shang-Bin Feng et al.· 0 citations
An extensive evaluation of automatic harness evolution for LLM agents is conducted, comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluating evolved harnesses on held-out tasks to assess whether the discovered improvements gen...
Yike Wang, Huaisheng Zhu, Zhengyu Hu et al.· arXiv.org· 18 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.