Jun 2026
Stage-Transition Dense Reward Modeling for Reinforcement Learning
Experiments show that STDR consistently improves sample efficiency and success rates over multiple baselines, and matches or surpasses handcrafted dense rewards on several challenging tasks, suggesting robustness to visual noise and better-calibrated reward assignment across settings.
Yang Yang, Bingjie Chen, Zihan Wang et al.
· arXiv.org · 0 citations