AgentRM: Enhancing Agent Generalization with Reward Modeling
This work finds that finetuning a reward model to guide the policy model is more robust than directly finetuning the policy model, and proposes AgentRM, a generalizable reward model, to guide the policy model for effective test-time search.
Yu Xia, Jing-Ru Fan, Weize Chen et al.
· Annual Meeting of the Associ... · 26 citations
· ⚡3