This work forms a post-failure decision as recovery routing over heterogeneous actions and train a supervised router from execution rollouts to make the same router usable under changing budgets, and adds a Conformal Risk Control layer that selects a deployment-time cost penalty without retraining and provides marginal...
Qi-Jia He, Jia-Yi Cheng, Chen-Qian Le et al.· arXiv.org· 2 citations
Agent reinforcement learning (RL) increasingly runs through full execution harnesses, and a multi-harness recipe mixes two choices: exposing the policy to several harnesses, and comparing their rewards inside one relative-advantage group. We isolate the second choice in repository-level coding. From one Qwen3-8B superv...
Chen-Qian Le, Jia-Yi Cheng, Qi-Jia He et al.· 1 citation
This method generates candidate responses from the target policy, evaluates helpfulness, factuality, and conciseness with rubric-specialized evaluators, applies a process-critic correction, and retains only high-consensus desirable or undesirable examples.
Bounded unfiltered teacher continuations at learner-induced contexts improve over pure behavioral cloning at matched budgets and suggest that a few teacher steps, placed at learner-induced contexts, can be a more cost-efficient supervision allocation than longer or more heavily curated teacher completions.
Junze Ye, Jiayi Cheng, Miao Lu et al.· arXiv.org· 2 citations
PreDiff-LM preserves causal attention within the observed prompt while allowing full bidirectional attention within the masked target, position hybrid attention as a complementary mechanism for adapting pretrained causal backbones, while making explicit the remaining quality and inference-efficiency gaps to optimized A...