Skip to content

On-Policy Distillation Meets Off-Policy GRPO: Training Compact Instruction-Following Rerankers

Sep 2026 · 0 citations · 32 references
Computer Science

TL;DR

Control comparisons against offline pairwise RankNet KD and on-policy GKD show that neither changing the offline distillation objective nor moving teacher-distribution matching on-policy reproduces the performance of reward-based on-policy distillation over student-sampled rankings.

Abstract

Compact instruction-following rerankers are attractive for deployment, but conventional distillation pipelines typically train students by offline imitation of teacher outputs on a fixed set of examples, constraining supervision to the teacher's observed ranking space. We revisit reranker distillation through the lens of reinforcement learning. We propose a two-stage framework combining off-policy teacher optimization with on-policy student distillation. In Stage 1, a 4B teacher reranker is strengthened with off-policy GRPO using LLM-judge feedback on 88K instruction-following examples. In Stage 2, a compact 1B student samples rankings from its own policy and receives soft teacher-derived rewards on those rankings, coupling student exploration with knowledge transfer. Our strongest gains appear under distribution shift. On MAIR-11, the original 11-subset, 869-query evaluation, the proposed student reaches 0.7670 nDCG@6, outperforming offline listwise KD by +4.6 points. Controlled comparisons against offline pairwise RankNet KD and on-policy GKD show that neither changing the offline distillation objective nor moving teacher-distribution matching on-policy reproduces the performance of reward-based on-policy distillation over student-sampled rankings. The advantage persists on MAIR-Full: across all 126 tasks and 9,356 queries, the proposed method obtains the highest task-macro point estimates among the evaluated distillation variants, reaching 0.6808 nDCG@6 and 0.7865 MRR@6. It also exceeds two released 7B RL-trained rerankers on the comparable MAIR-11 evaluation, while the same Stage 2 training procedure consistently improves three architecturally distinct alternative student backbones. On the 9,861-query validation benchmark, the resulting 1B reranker achieves 0.7624 nDCG@6 while providing a favorable quality-efficiency tradeoff relative to larger alternatives.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning

Adaptive On-Policy Self-Distillation is proposed, which adapts what information the teacher receives and how strongly its feedback influences learning, and encodes each solution as a reasoning DAG, orders problems by the student's evolving capability, and reveals only the affordable subgraph and its next frontier as PI...

Jia-Cheng Du, Wei-Wei Xie, Tian-Yi Du et al. · 0 citations
#machine learning Preprint Oct 2026

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optim...

Karn Tiwari, V. Chordia, P. PrathoshA · 0 citations
#artificial intelligence Preprint Oct 2026

Transfer-Stratified On-Policy Distillation for RL-Improved Reasoning Teachers

Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO,...

Xiao-Yu Chen, Bo Shao, Tian-Gang Zhu et al. · 0 citations
#machine learning Preprint Sep 2026

Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers

On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors can move students beyond the teacher's effective supervision. We introduce Graph-Conditioned On-Policy Agent Distillation (GC-OPD), which enriches an off-the-shelf...

Xiao-Han Yi, Wen Luo, Ya-Ni Huang et al. · 0 citations
#machine learning Preprint Sep 2026

REVO: Rollout-Efficient Off-Policy Distillation via Variance-Guided Reuse

On-policy distillation (OPD) trains language models using dense token-level teacher supervision on student-generated trajectories. However, its reliance on frequently refreshed student rollouts often incurs substantial generation cost. We introduce REVO, an off-policy distillation framework that improves rollout effici...

Yu-Xiao Yang, Shang-Zhe Li, Tian-Run Yu et al. · 0 citations
#machine learning Preprint Sep 2026

CLOOPD: Closing the Learner Loop in On-Policy Distillation

On-policy distillation (OPD) pays twice for each fresh batch: the student generates trajectories and a stronger teacher scores them. Existing methods improve which trajectories are scored and how the teacher signal is constructed, but usually consume it with one actor update. We introduce CLOOPD, a closed-loop framewor...

Keye Zheng, Han-Yu Li, Zhan Cheng et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.