Skip to content
Preprint

DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models

Aug 2026 · 0 citations · 39 references
Computer Science

TL;DR

This work proposes DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms, and converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD.

Abstract

Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learning enables direct optimization of task-specific rewards beyond the original models, yet trajectory-level optimization may incur high-variance gradients and cross-task interference. On-policy distillation (OPD) offers dense and stable supervision on student rollouts, but conventional teacher matching remains imitation-based. We propose DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms. Our DreOPD converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD. It further uses a mildly degraded reference to strengthen the teacher-reference contrast, yielding a clearer extrapolation direction. Experiments on single- and multi-teacher settings show that DreOPD outperforms OPD and multi-task RL baselines in average performance, while surpassing specialized teachers on most metrics.

View source

Similar papers

Preprint Aug 2026

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

Self-OPD is introduced, a teacher-free OPD framework for flow matching models that turns the student's own self-exploration into step-wise supervision and outperforms prior RL and OPD methods without task-specific teachers.

Shiyi Zhang, Mu-Shui Liu, Yunze Tong et al. · 0 citations
Jul 2026

FlowCTS: On-policy Continuous Trajectory Supervision of Flow Models

Flow Continuous Trajectory Supervision (FlowCTS), which matches subsequent student and reference trajectories initialized from the same student-visited state to derive a temporally weighted velocity-matching upper bound and discretize it into practical objectives parameterized by the number of supervision steps.

Kaiyang Ye, Yuan Ge, Junxia Zhang et al. · 0 citations
Jul 2026

Contrastive Reinforced Policy Optimization via Privileged Self-Distillation

Contrastive Reinforced Policy Optimization (CRPO) is introduced, which reformulates agentic OPSD from a contrastive learning perspective, and conducts group-wise contrast to preserve reliable, fine-grained optimization signals.

Xingjian Wu, Junlin Liu, Xing-Chen Liu et al. · 1 citation
Preprint Aug 2026

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation

REOPD combines a token-level compatibility weight with a batch-level adaptive budget, yielding a token-wise coefficient $\lambda_{b,t}=1+\gamma_b q_t$ that preserves teacher alignment while selectively extrapolating along reliable teacher-reference directions, demonstrating effective fine-grained reliability adaptation across domains and teacher configurations.

Yang Sun, Li-Chao Ma, Houyuan Qin et al. · 1 citation
Preprint Aug 2026

SR-OPSD: Self-Referenced On-Policy Self-Distillation

Extensive experiments across scientific evaluation, mathematical reasoning, and coding generation tasks with multiple large language models show that SR-OPSD achieves the state-of-the-art or competitive performance across various settings.

Zhuo Sun, Entong Li, Yan-Long Zhao et al. · 0 citations
Preprint Sep 2026

CA-OPD: Confidence-Aware On-Policy Distillation for Structured Visual Prediction

Autoregressive vision language models unify heterogeneous perception tasks but are highly susceptible to compounding errors. On-policy distillation (OPD) bridges the training-inference mismatch by training students on their own rollouts. However, unreliable student predictions, especially early in training, can derail the trajectory and degrade the quality of teacher supervision. While recent interleaved distillation methods allow the teacher to verify and replace student tokens, they primarily rely on rigid ranking metrics rather than exact teacher confidence, and they overlook how intervention decisions can inform token-level supervision. To address this, we introduce Confidence-Aware On-Policy Distillation (CA-OPD), a framework that couples reliable rollout construction with adaptive supervision. CA-OPD utilizes teacher confidence to selectively correct unreliable student transitions, gradually transferring rollout control to the student via a strict-to-relaxed schedule. Crucially, CA-OPD aligns knowledge transfer with these intervention decisions: corrected positions receive direct cross-entropy supervision from the teacher's prediction, while retained positions benefit from the teacher's full predictive distribution. Evaluated in a multi-teacher setting for GUI grounding and optical character recognition, CA-OPD substantially improves the Qwen3.5-0.8B baseline across all six target benchmarks, including gains of $9.50$ points on ScreenSpot-Pro and $6.72$ points on OCRBench-v2 English. Controlled studies further show that the gains depend on intervention placement, progressive rollout control, and intervention-aligned supervision, rather than intervention frequency alone.

Meng-Hao Li, Lin-Jie Mu, Yin Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.