Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO
Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation ac...