Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation
GAD-RL is introduced, which adaptively regulates teacher supervision during joint post-training according to the student's current task performance and local distributions, and weights forward KL by the student's probability of the teacher's Top-1 token, moderating local auxiliary updates when student support for that...