Skip to content
Preprint

Scaling Reinforcement Learning for Diffusion Models via Velocity Matching

Aug 2026 · 0 citations · 76 references
Computer Science

TL;DR

This work proposes reward-based velocity matching (RVM), a simple trajectory-free update that acts directly on the velocity field and provides a general framework that recovers recent fine-tuning methods, including RAM and DiffusionNFT, as special cases.

Abstract

Reward fine-tuning is becoming an important tool for adapting diffusion models to human preferences and task-specific objectives, but existing methods largely inherit policy-gradient machinery from large language models. Unlike autoregressive models, diffusion models do not provide tractable likelihoods for generated samples. As a result, current approaches either construct trajectory likelihoods from stochastic denoising transitions or approximate endpoint likelihoods with evidence lower bound, introducing additional computation and algorithmic complexity. We demonstrate that this likelihood-based machinery is not necessary for effective diffusion reward fine-tuning. We propose reward-based velocity matching (RVM), a simple trajectory-free update that acts directly on the velocity field. RVM reinforces directions associated with high-reward generations, suppresses those with low reward, and involves an optional anchor term controlling drift from a reference velocity. Notably, it provides a general framework that recovers recent fine-tuning methods, including RAM and DiffusionNFT, as special cases. Across various large-scale diffusion models reward fine-tuning tasks, RVM is competitive with or outperforms trajectory-based policy-gradient methods under substantially reduced training cost. We further find that, once the velocity update is simplified, the particular loss variant matters less than reward and anchor design. For video generation, standard preference rewards can favor visually clean but nearly static outputs; introducing a new dynamic-tracking reward that substantially improve motions while improving overall VBench performance. These results suggest that scalable reward fine-tuning for diffusion models is better posed in the native velocity representation than as likelihood-based policy optimization.

View source

Similar papers

Preprint Aug 2026

Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View

The empirical gap between these method families is identified as a variance-reduction effect rather than a difference in RL principle, and a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones is proposed.

Yi-Xian Xu, Yuanrui Zhang, Shengjie Luo et al. · 1 citation
Preprint Aug 2026

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models

Stage-Guided Per-Step Optimization (SGPO) is proposed for diffusion models, which jointly leverages signal-to-noise ratio and semantic changes to identify generation stages and adaptively assign stage-specific objectives.

Ren-Ye Yan, Ji-Kang Cheng, You Wu et al. · 0 citations
Preprint Aug 2026

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model

PAST is proposed, which provides differentiated rewards while adaptively regulating training episode length by jointly perceiving denoising progress and prompt difficulty and establishes a dual adaptive coordination mechanism that balances the extrinsic and intrinsic rewards.

Ren-Ye Yan, Ji-Kang Cheng, You Wu et al. · 0 citations
Preprint Aug 2026

A Mean-Field Framework for Inference-Time Distributional Control of Diffusion Models

Diffusion models are increasingly used as controllable samplers, whose generations can be steered at inference time according to a chosen reward function. While such rewards are typically defined on individual samples, for many applications it is desirable to steer according to distribution-level rewards, for example to calibrate with population-level information or to encourage diversity. In both cases, simply incorporating the reward gradient into the dynamics, while often effective, comes with few theoretical guarantees on the sampled distribution. For pointwise rewards, recent work has therefore sought to develop a principled framework for targeting a prescribed tilted distribution using particle reweighting. However, an analogous theoretically-grounded approach for distributional rewards is currently lacking. In this work, we formulate inference-time distributional control as targeting a tilted measure under a mean-field framework, and derive a weighted interacting particle scheme to target it in a principled manner. Our framework recovers pointwise-reward steering as a special case, while providing a theoretical foundation for existing batch-level steering methods. Empirically, we verify that the procedure correctly targets the prescribed distribution in tractable low-dimensional settings, and investigate its behaviour in higher-dimensional protein conformation tasks.

Samuel Howard, N. Nüsken · 0 citations
Preprint Aug 2026

Latent Reward Registers for Diffusion Preference Alignment

This work proposes Latent Reward Registers, a mechanism that estimates terminal preference directly from intermediate noisy latents, and achieves significant reward improvement with a favorable reward-quality balance against training-free baselines.

Yuanshen Guan, Zipeng Feng, Chengru Song et al. · 0 citations
Preprint Aug 2026

On-Policy Self-Distillation in Diffusion Models

The results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.

Weina Zhou, Xiongwei Zhu, Lingdong Kong et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.