Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
This work establishes a novel theoretical analysis: DDPO is an implicit form of score/flow matching with noisy targets, which increases variance and slows convergence, and introduces Advantage Weighted Matching (AWM), a policy-gradient method for diffusion.