Skip to content

Author

Guoping Pan

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jun 2026

Stage-Transition Dense Reward Modeling for Reinforcement Learning

Experiments show that STDR consistently improves sample efficiency and success rates over multiple baselines, and matches or surpasses handcrafted dense rewards on several challenging tasks, suggesting robustness to visual noise and better-calibrated reward assignment across settings.

Yang Yang, Bingjie Chen, Zihan Wang et al. · 0 citations