The concept of in-context self-evolution is formalized and VALVE, a validated-gated framework for long-horizon skill optimization is introduced, which establishes finite convergence, provides theoretical guarantees for future-task gain and drawdown, and derive the validation and evaluation holdout sizes required for a...
Yi-Fan Wang, Hao Cheng, Xiao-Min Li et al.· 0 citations
This work introduces ExecCritic, combining a test--verify--revise scaffold with a role-specific reinforcement learning recipe for training agents within it, separating test construction from source-code repair.
Leitian Tao, Bao-Lin Peng, Hao-Rui Wang et al.· 0 citations
This work compares six tool architectures that hold the underlying information and actions similar while varying how they are organized and exposed to the model, across three actors and a total of 11,700 trajectories to show that, even when tools provide similar capabilities, tool architecture changes agent behavior.
Xiangzhe Xu, H. Saghir, Qian-Hui Wu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.