The concept of in-context self-evolution is formalized and VALVE, a validated-gated framework for long-horizon skill optimization is introduced, which establishes finite convergence, provides theoretical guarantees for future-task gain and drawdown, and derive the validation and evaluation holdout sizes required for a...
Yi-Fan Wang, Hao Cheng, Xiao-Min Li et al.· 0 citations
This work introduces ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it, separating test construction from source-code repair.
Leitian Tao, Bao-Lin Peng, Hao-Rui Wang et al.· 0 citations
On long-horizon complex search, TRACE substantially improves base-model tool-use ability using pure RL, without a cold-start supervised fine-tuning stage, an agentic mid-training stage, or training on live-web data.
Leitian Tao, Baolin Peng, Wenlin Yao et al.· 4 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.