Skip to content

TISD: On-Policy Self-Distillation with Trajectory Intervention

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

A simple branch-regenerate-distill algorithm, Trajectory-Intervention Self-Distillation (TISD), which forces a teacher-selected branch action, returns suffix generation to the student, and distills the full trajectory under the privileged-context-conditioned teacher, support teacher-guided branching as a way to expose useful successor contexts for self-distillation.

Abstract

On-policy self-distillation (OPSD) provides dense teacher targets, but evaluates them only along student-sampled rollouts. When the privileged teacher favors an alternative action at a visited prefix, OPSD can provide a target for the branch decision but cannot supervise the successor contexts induced by that action unless the student samples it. This creates a training-time data-collection bottleneck and suggests a different role for teacher-student disagreement: proposing a trajectory branch rather than identifying a sufficient local repair. Our diagnostic framework using controlled token interventions reveals that a teacher-preferred token at peak disagreement can improve student continuation success, while its local corrective value is limited. Motivated by this finding, we introduce a simple branch-regenerate-distill algorithm, Trajectory-Intervention Self-Distillation (TISD). TISD forces a teacher-selected branch action, returns suffix generation to the student, and distills the full trajectory under the privileged-context-conditioned teacher. Across the coding models, TISD improves average Avg@4 over SDPO by 1.2 percentage points. Across the science domains, it improves average Avg@128 by 0.8 points under an equal-step budget and by 0.3 points under an equal-time budget. These results support teacher-guided branching as a way to expose useful successor contexts for self-distillation.

View source

Similar papers

#machine learning Preprint Sep 2026

Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factor...

Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni · 0 citations
Preprint Sep 2026

What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation

On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We fi...

Zi-Zhuo Lin, Quan-Ling Liu, Yi Yang et al. · 0 citations
#machine learning Preprint Sep 2026

Reward-Aligned Reweighting for On-Policy Distillation

On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however,...

Hao Xu, Junwei Su, Lan-Song Diao et al. · 0 citations
#machine learning Preprint Aug 2026

SR-OPSD: Self-Referenced On-Policy Self-Distillation

Extensive experiments across scientific evaluation, mathematical reasoning, and coding generation tasks with multiple large language models show that SR-OPSD achieves the state-of-the-art or competitive performance across various settings.

Zhuo Sun, Entong Li, Yan-Long Zhao et al. · 0 citations
Preprint Aug 2026

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

Self-OPD is introduced, a teacher-free OPD framework for flow matching models that turns the student's own self-exploration into step-wise supervision and outperforms prior RL and OPD methods without task-specific teachers.

Shi-Yi Zhang, Mu-Shui Liu, Yunze Tong et al. · 1 citation
Preprint Aug 2026

PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation

Privileged Adaptation from Student Trajectories (PAST), which treats each completed student trajectory as additional privileged information for the OPSD teacher while leaving the student's distillation prefixes unchanged, and improves the Avg@12 macro average over Vanilla OPSD by 5.6 percentage points.

Yang-Yang Feng, Zhuoyan Feng, Jun-Lan Chen · 5 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.