Preprint
Jul 2026
Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards
This work establishes the first non-vacuous generalization bounds for parameter-efficient RLVR fine-tuning at the billion-parameter scale, and proposes the Progressive RLVR framework, which integrates RLVR with on-policy distillation, TinyLoRA, and model quantization.
Yuxuan Zhu, Rohan Alur, Daniel Kang
· 0 citations