Skip to content
Preprint

Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation

Aug 2026 · 0 citations · 18 references
Computer Science

TL;DR

Findings establish recoverability as an outcome-grounded decision variable for selective supervision in OPD and show that retaining teacher-correctable prefixes provides the largest individual contribution.

Abstract

On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous prefix remains correctable. We formulate this decision as counterfactual recoverability and replay each error state through budget-matched teacher-continuation and rollback branches. Based on their relative success, states are categorized as recoverable, irreversible-but-avoidable, or ambiguous, and these labels guide whether training retains, rolls back, or conventionally supervises the corresponding trajectory. On AIME branch diagnostics, the mean continuation-minus-rollback effect is 0.185 for recoverable states and -1.000 for irreversible-but-avoidable states, demonstrating opposite intervention preferences. A branch-derived recoverability proxy achieves an AUC of 1.000, substantially outperforming divergence alone at 0.392. Across frozen evaluations, recoverability-aware control achieves the strongest recorded performance, reaching 0.578 success on held-out AIME2025 compared with 0.517 for the best baseline. It also improves AIME2024-2025 average@32 from 0.2656 to 0.3125 and GPQA-Diamond average@32 from 0.2702 to 0.3070. Component ablations further show that retaining teacher-correctable prefixes provides the largest individual contribution. These findings establish recoverability as an outcome-grounded decision variable for selective supervision in OPD.

View source

Similar papers

#machine learning Preprint Sep 2026

Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factor...

Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni · 0 citations
#artificial intelligence Preprint Sep 2026

Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives

On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student m...

Qi-Wei Di, Xu-Heng Li, Kaixuan Ji et al. · 0 citations
#natural language process... Preprint Aug 2026

CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation

On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this limitation by allocating supervision non-uniformly across response tokens according to their estimate...

Enhan Li, Jun-Hao He, Hong-Yang Du · 4 citations
Book Open access Aug 2026

Mining Point-of-No-Return Boundaries in Constrained Dynamical Systems via Counterfactual Auditing

Safety-critical failures in constrained dynamical environments are often detected only after a violation occurs, while outcome metrics (e.g., success rate, time-to-failure) conflate structural inevitability with decision-induced errors. We formulate failure diagnosis as recoverability boundary discovery and define the...

Jia Liu, Jiaxin Luo, Lejun Ai et al. · 0 citations
Preprint Aug 2026

Revelation Control

Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equ...

Qin-You Wang · 1 citation
Preprint Sep 2026

When Is Inaction a Mistake? Continuation-Aware Auditing of PPO Trading Policies

A four-stage audit for frozen proximal policy optimization policies without retraining examines deployment occupancy, matches current information, tests isolated deviations under incumbent continuation, and evaluates repeated deployment of observation-based alternatives.

Xing-Fei Zeng, Xin Zhong, Nan-Ting Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.