Reinforcement learning has substantially improved large language model (LLM) agents in verifiable domains, but remains difficult to apply to open-ended agent tasks, where solutions are diverse and reliable scalar rewards are hard to obtain. Recent pairwise evaluation methods alleviate reward discrimination collapse by...
Qiang Zhang, Rui-Xue Ding, Fanrui Zhang et al.· 1 citation
ARISE-RL, a novel full-cycle self-evolution framework that couples a task/rubric Generator and a reasoning Solver through rubric-mediated co-evolution, is proposed and ECR-Bench, an expert-calibrated rubric benchmark suite covering single-tool deep research and multi-tool travel planning is presented.
Fanrui Zhang, Rui-Xue Ding, Qiang Zhang et al.· 0 citations
Experimental results show that although current agents can reliably uncover the problems exposed by alerts, they struggle to proactively investigate the disk for silent intrusions and to produce comprehensive, verified remediation plans, with no model achieving complete detection and remediation on any single range.
Le-Han Wang, Boli Chen, Ruixue Ding et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.