Skip to content

SCOPE-RL: Optimizing Reasoning Paths Before and After Success

Jul 2026 · arXiv.org · Vol abs/2607.11506 · 0 citations · 28 references
Computer Science

TL;DR

SCOPE-RL improves average accuracy by up to 11.2 pp and reduces reasoning tokens by up to 27.1% over outcome-only GRPO, indicating that reward-signal densification is complementary to policy-update-level RLVR advances.

Abstract

Reinforcement learning with verifiable rewards (RLVR) optimizes LLMs using sparse verifiable final-answer rewards. This sparse anchor reliably verifies whether a trajectory succeeds but provides no direct feedback on the reasoning path that produced it. Before success, prerequisite progress on hard problems receives no reward signal; after success, outcome rewards cannot distinguish well-organized correct trajectories from redundant or locally flawed ones. We introduce SCOPE-RL (Scaffolded Chain Optimization with Process Efficiency), a two-stage framework that densifies this anchor while retaining the GRPO update: Adaptive Scaffolded RL adds prefix-decomposed verifiable rewards on answer-hidden sub-question chains before success, and Quality-Aware Process RL applies correctness-gated process-shape rewards to refine correct trajectories after success. An expert-validated Step-Quality Evaluation Protocol evaluates useful-step density, error localization, and token efficiency beyond final-answer accuracy. On Qwen3-8B-Instruct trained on DAPO-Math and Big-Math, SCOPE-RL improves average accuracy by up to 11.2 pp and reduces reasoning tokens by up to 27.1% over outcome-only GRPO; the gains hold under GSPO and on Qwen3-0.6B-Instruct, indicating that reward-signal densification is complementary to policy-update-level RLVR advances. Code and data are available at https://github.com/tokencraft-lab/SCOPE-RL.

View source

Similar papers

#natural language process... Preprint Sep 2026

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying

Reinforcement learning (RL) has become one of the primary paradigms for reasoning enhancement of large language models (LLMs). In particular, Group Relative Policy Optimization (GRPO) and related algorithms have demonstrated strong performance with outcome-level rewards. However, these methods depend solely on the fina...

Shi-Qi Yan, Chao-Hong Tan, Qian Chen et al. · 0 citations
Preprint Aug 2026

StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning

This work introduces StructReward, a compute-efficient framework that provides dense reinforcement signals through structured step-level reward alignment and substantially reduces the computational overhead of multimodal reinforcement learning.

Yifan Li, Ruxi Sun, Tong-Zhou Zhao · 0 citations
#machine learning Preprint Sep 2026

Gradients Know What Outcomes Don't: Unlocking Reinforcement Learning for LLM Reasoning with Gradient-Aligned Rewards

Gradient-Aligned Reward (GAR), which operates in the policy's own gradient space: truncated backpropagation through the output projection layer extracts a compact gradient vector for each rollout, and cosine similarity with an expert-anchor gradient yields a dense, reasoning-aware reward with less than 9% wall-clock ov...

Le-Qi Zheng, Jin-Bo Su, Fang Niu et al. · 2 citations
#machine learning Preprint Jul 2026

Verifier-Induced Support Reshaping in On-Policy Optimization

It is shown that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce, and effective rewardable support is defined as successful trajectories reachable within a fixed rollout budget.

Shaohang Wei, Z.Y. Su, Feifan Song et al. · 0 citations
#natural language process... Preprint Sep 2026

From Base Rollouts to RL Reasoning: A Budgeted Search Perspective

Results support a qualified internalized-search reading: under the recipe the authors test, much of the measured RL gain corresponds to a change in sampling efficiency toward operating points the base model can already reach under search.

Wen-He Sun, Cun-Xiang Wang, Zi-Jun Yao et al. · 0 citations
#artificial intelligence Preprint Aug 2026

On-policy Distillation with Verifiable Reward

This work reformulates the implicit reward of sampled-token OPD based on trajectory correctness, then applies a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards, making it readily combinable with any policy gradient algorithm, such as...

Wen-Ze Lin, Jia-Le Zhao, Xi-Tai Jiang et al. · 9 citations · ⚡4

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.