Skip to content

RLVP: Penalize the Path, Reward the Outcome

Jul 2026 · arXiv.org · Vol abs/2607.07435 · 0 citations · 52 references
Computer Science

TL;DR

The resulting recipe"penalize the path, reward the outcome" achieves high task success with near-zero violations, where outcome-only training violates constraints on nearly every episode.

Abstract

Agents acting on our behalf in the real world (e.g. placing phone calls) must learn online from costly, often irreversible interactions rather than cheap simulator steps. Two things follow. First, deployability depends on the path, not only the outcome. An agent must respect outcome-neutral constraints such as not repeatedly calling an unresponsive user, respecting business hours, or completing required authentication constraints that outcome-based rewards cannot express, since violating them frequently improves apparent success. Second, because each interaction is expensive, the agent must learn efficiently from very few examples. Reinforcement learning from verifiable rewards (RLVR) is blind to both challenges: it optimizes solely on the outcome and wastes expensive rollouts on all-fail groups where group-relative advantage collapses to zero. Attempts to densify supervision by rewarding progress target the hard-to-verify direction. In contrast, real agentic environments can cheaply detect bad moves. Since group-relative advantage is equivalent to within-group variance, a dense signal helps only when it supplies variance the outcome lacks. A verifiable penalty on the path meets this condition reliably, while a progress potential helps only where partial progress is reachable. The resulting recipe"penalize the path, reward the outcome"achieves high task success with near-zero violations, where outcome-only training violates constraints on nearly every episode. We provide four design rules for effective penalties, including avoidance of the inaction trap that arises when a penalty is used in isolation.

View source

Similar papers

#machine learning Preprint Jul 2026

Verifier-Induced Support Reshaping in On-Policy Optimization

It is shown that on-policy reinforcement learning with verifiable rewards (RLVR) can improve the current objective while making successful behaviors for later objectives too rare to sample and reinforce, and effective rewardable support is defined as successful trajectories reachable within a fixed rollout budget.

Shaohang Wei, Z.Y. Su, Feifan Song et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Spurious Advantage Hidden in GRPO

SignBALANCE is proposed, whose magnitude is composition-free: it keeps the verifier sign, uses a global scale, and restores zero-mean balance via a stop-gradient per-class rescaling.

Jiamian Wang, Samyadeep Basu, Koustava Goswami et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Explore More, Drift Less: Outcome-Only Reinforcement Learning Can Suffice for Long-Horizon Interactive Agents

CANOPY (Coverage-ANchored On-PolicY RL), a minimalist protocol attacking both Signal starvation and policy drift, internalizes long-horizon capability directly into a small open model; the complete training stack is planned to be released at https://github.com/AlibabaResearch/SignalCoverageRL.

Li-Ming Pu, Xiao-Xiao Li, Yi-Fu Liu et al. · 0 citations
Preprint Aug 2026

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), which treats teacher reliance as capability-dependent and uses imitation only where group-relative rewards are uninformative, obtains the best result in all four scientific domains.

Yubo Zhang, Xin-Hong Ma, Zezhong Tan et al. · 1 citation

Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information

Off-Context GRPO (OC-GRPO), a minimally modified variant of GRPO that uses guided rollouts but applies an importance-corrected objective to steer the update back toward the original unguided objective, avoiding the mismatch that destabilizes uncorrected guided training.

Priyank Agrawal, Ankur Samanta, S. Ghasemlou et al. · 1 citation
Jul 2026

Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks

Progress-conditioned Group Policy Optimization is proposed, which uses first-visit observation coverage only when all samples in a group receive zero outcome reward, and consistently improves over group-based baselines, with particularly large gains on hard tasks.

Kaibing Yang, Guangfeng Cai, Sheng-Tian Yang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.