This work introduces ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it, separating test construction from source-code repair.
Leitian Tao, Bao-Lin Peng, Hao-Rui Wang et al.· 0 citations
On long-horizon complex search, TRACE substantially improves base-model tool-use ability using pure RL, without a cold-start supervised fine-tuning stage, an agentic mid-training stage, or training on live-web data.
Leitian Tao, Baolin Peng, Wenlin Yao et al.· 4 citations
Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems, and the 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form.
B. An, B. Li, B. Wang et al.· 3 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.