Preprint
Jul 2026
dOPSD: On-Policy Self-Distillation for Diffusion Language Models
dOPSD derives the teacher's privilege directly from the student's own denoising trajectory, evaluating masked positions using later, more-decoded steps of that same trajectory rather than an external label, so the teacher's advantage emerges from the model's own decoding process.
P. Dat, Qi Li, Xinchao Wang
· 0 citations