On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers
Control comparisons against offline pairwise RankNet KD and on-policy GKD show that neither changing the offline distillation objective nor moving teacher-distribution matching on-policy reproduces the performance of reward-based on-policy distillation over student-sampled rankings.
Vignesh Prabhakar, Jiagi Pan, Anil Babu Ankisettipalli
· 0 citations