Preprint
Aug 2026
DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models
This work proposes DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms, and converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD.
Ming-Hung Lin, Chengfei Cai, Lin Xu et al.
· 0 citations