This work finds that finetuning a reward model to guide the policy model is more robust than directly finetuning the policy model, and proposes AgentRM, a generalizable reward model, to guide the policy model for effective test-time search.
Yu Xia, Jing-Ru Fan, Weize Chen et al.· Annual Meeting of the Associ...· 26 citations· ⚡3
StatFormBench is introduced, a benchmark built from five cross-domain statistics textbooks and a data science case library, covering diverse problem types, data representations, and scenario styles, and no model performs consistently best across the two subtasks.
Chen Wang, Jun-Zhe Zhao, Xin Cong et al.· 0 citations
Analyzing a large corpus of publicly released post-training trajectories, it is found that across different tasks, the agent's training strategy is locked in at the very beginning, and the entire remaining budget is spent on local adjustments within the selected strategy.
J. Lim, Xinyuan Huang, Hao Peng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.