Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control
POGP is introduced, a framework that learns a prefix value function at every intermediate denoising step through a Bellman-style recursion over the denoising chain, and indicates that supervising intermediate denoising steps is useful not only for adaptive early stopping, but also as an auxiliary objective that improve...