Large reasoning models (LRMs) have shown exceptional performance in complex tasks such as mathematics and coding. In the field of machine translation (MT), reinforcement learning (RL) has been utilized to enhance the quality of translations. However, traditional RL approaches rely heavily on the base model’s inherent...
Zengkui Sun, Jia-Li Zeng, Jiaan Wang et al.· Transactions of the Associat...· 0 citations
EvoBrowseComp is introduced, an evolving benchmark of 400 English and 400 Chinese contamination-free complex questions synthesized via live-web traversal that establishes a scalable paradigm for auto-updatable, high-difficulty benchmarking that keeps pace with both evolving world knowledge and advancing agent capabilit...
AlgoWorlds is introduced, a benchmark that transforms formally specified combinatorial optimization problems into partially observed decision environments with verifiable global optima, and seven leading LLMs are evaluated, including Claude Opus 4.8 and GPT-5.6 Sol.
Zi-Xiang Xu, Jiaan Wang, Fanfei Meng· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.