Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-s...
Qi-Jia He, Rui-Nan Jin, Jun Luo et al.· 0 citations
Beam-search-based test-time methods provide an effective way to improve large language model (LLM) performance on long-horizon generation by pruning invalid reasoning paths early, leading to significantly improved reasoning efficiency and more favorable test-time cost scaling. Despite strong empirical success, the theo...
Qi-Jia He, Yu Huang, Yuan Cheng et al.· 0 citations
This work forms a post-failure decision as recovery routing over heterogeneous actions and train a supervised router from execution rollouts to make the same router usable under changing budgets, and adds a Conformal Risk Control layer that selects a deployment-time cost penalty without retraining and provides marginal...
Qi-Jia He, Jia-Yi Cheng, Chen-Qian Le et al.· arXiv.org· 2 citations
Agent reinforcement learning (RL) increasingly runs through full execution harnesses, and a multi-harness recipe mixes two choices: exposing the policy to several harnesses, and comparing their rewards inside one relative-advantage group. We isolate the second choice in repository-level coding. From one Qwen3-8B superv...
Chen-Qian Le, Jia-Yi Cheng, Qi-Jia He et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.