Skip to content
Preprint

VERPO: Verified Evidence Regularized Policy Optimization

Sep 2026 · 0 citations · 37 references
Computer Science

Abstract

Verifiable rewards improve language models through reliable task-level feedback, but methods based on Group Relative Policy Optimization (GRPO) apply a sequence-level advantage uniformly across all tokens. This coarse credit assignment reinforces or penalizes entire responses without identifying which local decisions to preserve, reinforce, or revise. Conversely, evidence-conditioned self-distillation provides denser token-level supervision, yet teacher imitation can transfer stylistic artifacts and miscalibrated confidence that destabilize training when misaligned with task success. We introduce VERPO, which converts evidence-conditioned guidance into reward-aligned token-level credit assignment while retaining the outcome objective. VERPO decomposes teacher guidance into an evidence-free reference term and signed, evidence-induced corrections at each token. A stopped controller combines selective acceptance, token-wise localization, and cost-aware scaling by balancing alignment with the local GRPO update direction against Fisher movement cost. Furthermore, we introduce Fisher Evidence Contrast (FEC), which attenuates nuisance shifts along an estimated evidence-presence direction through a regularized projection. Across five scientific reasoning and tool-use tasks, VERPO prevents optimization collapse and consistently achieves the highest multi-task average across model backbones, yielding marked improvements particularly on smaller models over strong baselines. Qualitative diagnostics confirm that token acceptance selectively targets reasoning bottlenecks consistent with local reward alignment and Fisher movement cost.

View source

Similar papers

#machine learning Preprint Sep 2026

Reward-Aligned Reweighting for On-Policy Distillation

On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however,...

Hao Xu, Junwei Su, Lan-Song Diao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UniOPSD: Unifying Outcome and Hindsight Feedback for Agentic Reinforcement Learning

Reinforcement learning has become an effective approach to training language model agents, but sparse and delayed outcome rewards provide limited guidance for credit assignment across long interaction sequences. Recent work on on-policy self-distillation (OPSD) offers complementary supervision by evaluating a policy's...

Zeng-Huang Fu, Zhao-Yang Li, Qiu-Yuan Ai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PACT: From Credit Assignment to Critic Alignment

Policy Aligned Critic Training (PACT), which adopts an Actor-then-Critic update order to apply importance sampling correction to critic training and better align the critic with the updated policy, is developed.

Jia-Yan Fu, Hang Xu, Yong Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

Group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance, and demonstrates that group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance.

Zhu Zhang, Ji-Xun Wang, Xiao-An Xu et al. · 2 citations
#artificial intelligence Preprint Sep 2026

Beyond Token-Local Imitation: Reward-Compatible Temporal Credit Assignment for On-Policy Distillation

On-policy distillation (OPD) has emerged as an effective approach for large language model post-training, yet existing objectives face a trade-off between objective fidelity and optimization stability. Token-level OPD provides stable but local supervision, whereas sequence-level OPD captures future credit at the cost o...

Shi-Qi Liu, Ze-Yu He, Le-Tian Tao et al. · 2 citations
Preprint Aug 2026

Contrastive Branch Policy Optimization

Requiring only outcome rewards and no process-level annotation, CBPO provides a practical solution for fine-grained credit assignment in tool-integrated agent training and consistently outperforms state-of-the-art policy-optimization and branch-based methods.

Ying Wang, Changlin Qiu, Bang Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.