Skip to content

Author

Jianfei Zhao

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective

ConSPO is a Contrastive Sequence-level Policy Optimization method that uses length-normalized sequence log-probabilities as rollout scores and contrasts verified positive rollouts against negative distractors within the same group, and outperforms strong baselines on challenging reasoning benchmarks.

Feng Zhang, Xin-Hong Ma, Zi-Qiang Dong et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.