Skip to content
Preprint

On-Policy Delta Distillation for Multilingual Math Reasoning

Aug 2026 · 0 citations · 15 references
Computer Science

TL;DR

It is found that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.

Abstract

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delta Distillation (OPD$^2$), for mathematical reasoning in English, Korean, and Japanese. OPD$^2$ improves OPD by using the probability gap between a post-trained teacher and its base model as the learning signal. Experiments with Qwen3 show that OPD$^2$ consistently outperforms the original OPD, with particularly strong improvements in Korean and Japanese, and generally narrows the English-Korean performance gap. We further find that English-only OPD can also increase performance for Korean and Japanese, but often shifts the responses toward English, highlighting the importance of multilingual data to preserving target-language responses.

View source

Similar papers

Preprint Aug 2026

WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training

WDL-OPD is introduced, a mixture-constrained co-training method with two trainable policies that shows that freezing the auxiliary recovers an anchor-plus-contrast proxy target closely related to OPD$^2$ and W2S-OPD, whereas joint training creates branch-level degrees of freedom that a static delta cannot express.

Zehao Chen, Gong-Xun Li, Tianxiang Ai et al. · 0 citations
Jul 2026

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

Contrastive Reinforced Policy Optimization (CRPO) is introduced, which reformulates agentic OPSD from a contrastive learning perspective, and conducts group-wise contrast to preserve reliable, fine-grained optimization signals.

Xingjian Wu, Junlin Liu, Xing-Chen Liu et al. · 1 citation
Preprint Aug 2026

Tail-Aware Top-$k$ On-Policy Distillation

Tail-Aware Top-$k$ OPD is proposed, a novel distillation method that restores the missing tail probability signal and better aligns the student's next-token distribution with the teacher's, preventing the increase in tail probability and entropy caused by top-$k$ normalization.

Huipeng Huang, Hong-Xin Wei · 1 citation
Jul 2026

Flux-OPD: On-Policy Distillation with Evolving Contexts

Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.

Yu-Ran Wang, Zekun Wang, Bohan Zeng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Sequential Beats Joint: On the Interplay between On-Policy Distillation and RLVR

Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's dense token-level supervision to complement the sparse RL reward, fusing the two signals within a single step: either as a \emph{weighted-additive combination} or a \emph{teacher-modulated rescaling} of the RL advantage. In this paper, we show that a simple two-stage scheme, OPD-then-RL, consistently outperforms pure OPD, pure RLVR, and all such joint baselines across logic and math reasoning benchmarks. Beyond the empirical results, we further provide a systematic understanding of this through pass@$k$ behavior, learning dynamics, and parameter updates, yielding a consistent explanation: OPD expands the student's coverage of teacher-supported solutions and RL sharpens within that support, while jointly optimizing the two signals causes them to interfere. To provide a practical recipe, we find that the OPD validation score is the key signal for when to switch to RL, and that OPD is a better cold start for RL than SFT. Together, our results establish OPD-then-RL as a simple yet strong way to combine the two methods, turning two entangled signals into complementary stages.

Bo-Yang Li, Bingsen Chen, Cheng-Hao Yang et al. · 1 citation
Jul 2026

β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation

This work introduces $\beta$-OPSD and derives its optimal policy as a geometric interpolation between the reference policy and the privileged teacher, and provides a principled route from self-distillation to policy optimization and back without sacrificing the efficiency that makes OPSD practical.

Jiawei Xu, Ming-Hui Liu, Juzheng Zhang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.