ReCo, a reweighting method that improves Pass@k for large values of k and is comparable to GRPO for small values of k, and replaces the token-level importance ratio with a variance-based ratio that gives larger update scale to non-saturated decision points.
Abstract
Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that GRPO can reduce the base model's reasoning capacity and underperform it in Pass@k when k is large, indicating reduced coverage of reasoning paths. We find that this reduction is associated with GRPO concentrating on responses that the base model already generates with high probability. We trace this concentration to two mechanisms in the GRPO update. At the response level, high-probability responses dominate the group gradient through repeated occurrence. At the token level, GRPO's importance ratio scales gradients, further reinforcing tokens that become more likely under the current policy. We propose ReCo, a reweighting method that addresses both effects. Response contributions are normalized by their expected occurrence within the rollout group, and the token-level importance ratio is replaced with a variance-based ratio that gives larger update scale to non-saturated decision points where alternative token choices remain plausible. Across Qwen2.5-Math-1.5B/7B and Llama-3.1-8B-Instruct on five mathematical reasoning benchmarks, ReCo improves Pass@k for large values of k and is comparable to GRPO for small values of k.
Cue-GRPO improves AIME repeated-sampling performance and introduces a partition- conditioned rule that redistributes positive advantages accord- ing to cluster rarity, with Strategy Cues providing a low-overhead implementation for competition mathematics.
Gradient Uncertainty-Aware Policy Optimization is proposed, which models each group gradient as a random variable under a Bayesian formulation and estimates its probability distribution and derives gradient uncertainty using a Dirichlet-based formulation and uses it to calibrate the contribution of each group gradient...
Peizheng Guo, Jian-Qi Zhang, Xing-Yu Zhang et al.· 0 citations
SignBALANCE is proposed, whose magnitude is composition-free: it keeps the verifier sign, uses a global scale, and restores zero-mean balance via a stop-gradient per-class rescaling.
Jiamian Wang, Samyadeep Basu, Koustava Goswami et al.· 0 citations
I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), which treats teacher reliance as capability-dependent and uses imitation only where group-relative rewards are uninformative, obtains the best result in all four scientific domains.
Yubo Zhang, Xin-Hong Ma, Zezhong Tan et al.· 0 citations
Group-based reinforcement learning (RL) methods, such as GRPO and its variants, have become a leading paradigm for training reasoning and agentic large language models (LLMs). While their group-normalized advantage estimation is reliable at the response level, it becomes systematically biased at the step level, since c...
Xin-Cheng Yao, Haobo Fu, Wei-Ming Liu et al.· 0 citations
This work proposes two complementary strategies to improve the performance of value function RL: Privileged Value Functions (PVF) which provide an elegant mechanism to inject additional task-relevant token-level signal without biasing the policy objective; and TETHER, a baseline that adaptively interpolates between gro...
S. Venkatraman, Matthieu Dinot, Laurence Aitchison· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.