Skip to content

Recursive Self-Improvement via On-Policy Distillation for Reasoning

Sep 2026 · 0 citations · 39 references
Computer Science

TL;DR

DCE+SRCL outperforms OPSD across multiple model scales and four competition-level mathematics benchmarks, and its comprehensive evaluations show that DCE+SRCL outperforms OPSD across multiple model scales and four competition-level mathematics benchmarks.

Abstract

On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's next-token predictions. This provides dense, token-level supervision to the student. On-policy self-distillation (OPSD) eliminates the need for the external teacher. Specifically, a second frozen copy of the student model, now given the ground truth in its context, serves as the teacher. The student model only receives the problem and learns to mimic the privileged teacher model, while the teacher remains frozen throughout training. Previous work showed that freezing the teacher is useful for training stability, but we argue that this can prevent the teacher from incorporating the improvements learned by the student during training. Our primary contribution is to address this limitation with a recursive framework built around two complementary components. First, we let the privileged teacher co-evolve with the student so that revision learned in one round can guide the next, a process we refer to as Dynamic Co-Evolution (DCE). Second, because stronger revision can also make responses too verbose and self-critical, we additionally train on shorter, verified rewrites of the model's own on-policy responses. We call this complementary objective Self-Refined Concise Learning (SRCL). Overall, our comprehensive evaluations show that DCE+SRCL outperforms OPSD across multiple model scales and four competition-level mathematics benchmarks. Specifically, on Qwen3-8B, DCE+SRCL reaches 65.97% Average@12, outperforming OPSD by 35.62 percentage points while reducing mean output length by 7.80% relative to DCE alone.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Train Ahead, Distill Back: Bootstrapping On-Policy Self-Distillation for Large Language Models

On-policy self-distillation (OPSD) improves large language models by letting a self-teacher with privileged information provide dense token-level supervision on the model's own trajectories. Yet existing methods typically construct the self-teacher from the current, initial, or slowly averaged policy state, leaving the...

Zheng Zhang, Xin-Yue Tan, Lu Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning

On-policy self-distillation (OPSD) trains a question-only student with token-level feedback from a teacher given training-only privileged information (PI). OPSD therefore provides dense, on-policy supervision, and is free of a larger external teacher, but its effectiveness rests on how PI is designed and utilized. Our...

Jia-Cheng Du, Wei-Wei Xie, Tian-Yi Du et al. · 0 citations
#machine learning Preprint Sep 2026

Teach Yourself Where to Look: On-Policy Attention Self-Distillation for Reasoning

On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation, shows that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.

Safaeid Hossain Arib, Rabeya Akter, I. N. Swapnil et al. · 0 citations
Preprint Aug 2026

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

Self-OPD is introduced, a teacher-free OPD framework for flow matching models that turns the student's own self-exploration into step-wise supervision and outperforms prior RL and OPD methods without task-specific teachers.

Shi-Yi Zhang, Mu-Shui Liu, Yunze Tong et al. · 0 citations
Preprint Sep 2026

What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation

On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We fi...

Zi-Zhuo Lin, Quan-Ling Liu, Yi Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have v...

Zhen-Yu Wang, Tian-Ze Wang, Lin-Jun Zhang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.