Diffusion models faithfully reproduce their training distribution, but also inherit its imbalances and leave rare or under-represented modes hard to reach. A natural inference-time remedy is to sample from the high-temperature target $p^{(\gamma)}_0(x) \propto p_0(x)^{\gamma}$ for $0<\gamma<1$, which flattens dominant modes and lifts rare ones. However, naive score scaling while correctly reweighting modes also inflates the per-mode variance, breaking the reverse diffusion process and degrading sample quality. We introduce variance-corrective time shifting, a training-free fix that queries the network at a shifted timestep and scales the resulting score by $\gamma$, canceling the variance inflation while preserving the mode reweighting. The correction turns simple temperature sampling into a practical diversity knob for pretrained diffusion and flow-matching backbones with no retraining, and we demonstrate consistent gains at minimal cost to sample quality and condition fidelity across DiT, Stable Diffusion and Motion Diffusion models. We further show that the timing of the temperature intervention enables coarse-to-fine control: high-noise stages drive compositional diversity across modes, while low-noise stages drive local appearance variation under a fixed composition.
Peizhuo Li, Emre Aksan, A. Ichim et al.· 0 citations
We introduce stylized phase manifolds—a compact, interpretable latent representation that disentangles motion content (e.g. “jumping”, “walking”), the temporal structure (e.g. motion cycle frequency, gait timing), and style (i.e. how the motion is performed). Learned in an unsupervised manner and inherently low‐dimensional, the manifold offers intuitive and flexible editing. Building on this representation, we develop a diffusion‐based motion generator that enables fine‐grained control over semantic, temporal, and stylistic aspects of motion. To connect high‐level intent with low‐level motion, we treat the stylized manifold as an intermediate representation—a structured bridge between natural language and motion. By first mapping text into this manifold, our two‐stage pipeline improves the control over for text‐based motion generation, while producing high‐quality, diverse motion outputs.
Jingyuan Li, Peizhuo Li, A. Aristidou et al.· Computer graphics forum (Pri...· 0 citations