AgentRM: Enhancing Agent Generalization with Reward Modeling
This work finds that finetuning a reward model to guide the policy model is more robust than directly finetuning the policy model, and proposes AgentRM, a generalizable reward model, to guide the policy model for effective test-time search.