Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge.
Austin S. Wang, Zi-Heng Cheng, Le-Xing Ying
· 0 citations