Preprint
Aug 2026
Tail-Aware Top-$k$ On-Policy Distillation
Tail-Aware Top-$k$ OPD is proposed, a novel distillation method that restores the missing tail probability signal and better aligns the student's next-token distribution with the teacher's, preventing the increase in tail probability and entropy caused by top-$k$ normalization.
Huipeng Huang, Hong-Xin Wei
· 1 citation