Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
Switch Distillation is proposed, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy, which consistently outperforms existing distillation objectives across teacher sizes.