Skip to content
Preprint

Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control

Aug 2026 · 2 citations · 30 references
Computer Science

TL;DR

POGP is introduced, a framework that learns a prefix value function at every intermediate denoising step through a Bellman-style recursion over the denoising chain, and indicates that supervising intermediate denoising steps is useful not only for adaptive early stopping, but also as an auxiliary objective that improves the learned policy.

Abstract

Diffusion policies are a powerful policy class for continuous control, but their iterative denoising process creates a substantial computational bottleneck. Reducing this cost requires adapting the number of denoising steps to the difficulty of each action while preserving task performance. We introduce Prefix-Optimal Generative Policies (POGP), a framework that learns a prefix value function at every intermediate denoising step through a Bellman-style recursion over the denoising chain. The prefix value function serves two purposes: it provides an auxiliary training objective that encourages intermediate outputs to become high-quality actions, and it enables a test-time stopping rule that terminates denoising when additional steps are unlikely to produce meaningful improvement. Across four MuJoCo environments and comparisons with 12 baselines, POGP reduces the required number of denoising iterations by approximately 2.7-fold while retaining near-full task performance. Compared with state-of-the-art dynamic diffusion baselines, prefix training also improves final task performance by approximately 3.5%. These results indicate that supervising intermediate denoising steps is useful not only for adaptive early stopping, but also as an auxiliary objective that improves the learned policy.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Noisy-Space Policy Gradient for Diffusion Policies in Offline Reinforcement Learning

Diffusion policies offer a powerful and expressive parameterization for continuous control. Yet, their integration with reinforcement learning remains conceptually and algorithmically challenging. In this work, we address this gap by introducing a noisy-space action-value (Q-)function that assigns values to diffusion l...

Mahmoud Selim, Cristina Cipriani, K. H. Johansson · 0 citations
Preprint Sep 2026

DIA: Denoising Intermediate Advantage for Diffusion Policy Optimization

DIA is introduced, a policy-gradient method that learns a value function over partially denoised actions and uses it to construct a denoising level advantage for each step of the generative process, enabling it to discover more effective and efficient task-level strategies and subtask sequences that baseline methods fa...

Arjun Sohal, Yuchi Zhao, Miroslav Bogdanovic et al. · 0 citations
Preprint Aug 2026

Explore or Converge? Stage-Guided Per-Step Optimization for Diffusion Models

Stage-Guided Per-Step Optimization (SGPO) is proposed for diffusion models, which jointly leverages signal-to-noise ratio and semantic changes to identify generation stages and adaptively assign stage-specific objectives.

Ren-Ye Yan, Ji-Kang Cheng, You Wu et al. · 1 citation
Preprint Aug 2026

On-Policy Self-Distillation in Diffusion Models

The results support on-policy self-distillation as an efficient and analyzable approach to diffusion post-training by converting image-level reward guidance into explicit and continually refreshed intermediate supervision, thereby opening a path toward more efficient and diagnosable alignment.

Weina Zhou, Xiongwei Zhu, Ling-Dong Kong et al. · 4 citations
Preprint Aug 2026

PAST: Prompt-Adaptive Sampling Termination for Efficient Diffusion Model

PAST is proposed, which provides differentiated rewards while adaptively regulating training episode length by jointly perceiving denoising progress and prompt difficulty and establishes a dual adaptive coordination mechanism that balances the extrinsic and intrinsic rewards.

Ren-Ye Yan, Ji-Kang Cheng, You Wu et al. · 0 citations
#machine learning Preprint Aug 2026

Scaling Reinforcement Learning for Diffusion Models via Velocity Matching

This work proposes reward-based velocity matching (RVM), a simple trajectory-free update that acts directly on the velocity field and provides a general framework that recovers recent fine-tuning methods, including RAM and DiffusionNFT, as special cases.

Jaemoo Choi, Wei Guo, Yuchen Zhu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.