Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool responses offers limite...
Shuang Sun, Guo-Xin Chen, Fan-Zhen Meng et al.· 0 citations
Evo-Bench is the first benchmark designed to evaluate models'intrinsic harness-evolving capabilities across Search, Office, and General agent domains, and exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consis...
This work presents a unified black-box RL framework for stable and scalable optimization of general agents through complex harnesses, supporting unified training across heterogeneous execution systems.
Hua-Tong Song, Fei Bai, Ming Yang et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.