Skip to content

Author

Shilong Zhang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models

This work establishes a novel theoretical analysis: DDPO is an implicit form of score/flow matching with noisy targets, which increases variance and slows convergence, and introduces Advantage Weighted Matching (AWM), a policy-gradient method for diffusion.

Shuchen Xue, Chongjian Ge, Shilong Zhang et al. · 52 citations · ⚡11

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.