Equilibrium Forcing is introduced, a simplified framework for video denoising generative models without noise level conditioning that pioneers modular training- and inference-time designs for noise-unconditional generation that decouple learning the denoising field from sampling.
Abstract
Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampling schedules, limiting inference procedures from adapting to the data. We introduce Equilibrium Forcing (EqF), a simplified framework for video denoising generative models without noise level conditioning. EqF pioneers modular training- and inference-time designs for noise-unconditional generation that decouple learning the denoising field from sampling. This flexibility allows for inference-time algorithms that operate in a closed loop by adapting to feedback from the sample, improving video quality and consistency on challenging autoregressive video generation benchmarks. Extensive analysis elucidates exactly how removing the noise level conditioning enables EqF's data-dependent inference properties to surpass the performance of standard noise level-conditional denoising video methods.
Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion teacher and an online fake-score model. We instead learn the rollout distribution directly from reference videos, eliminating both score models during post-training. Our framework mi...
Chi Zhang, Yue-Yi Liu, Hao-Yan Shi et al.· 1 citation
This work reformulates the video diffusion sampling as a frame-indexed stochastic process over noise levels, and constructs a continuous training trajectory along which the sampling schedule progressively evolves from independent sampling to inference-consistent sampling.
Yue-Ting Zhu, Yue-Hao Song, Kai-Chen Zhang et al.· 1 citation
In-Context Forcing is introduced, a progressive autoregressive paradigm that utilizes contexts with decreasing noise levels that enables cross-frame parallel denoising, achieving substantial inference acceleration without sacrificing performance.
Lingxiao Yang, Liu Liu, Mo-Ran Li et al.· 2 citations
Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smooth...
Zhuo-Ran Zhao, Sheng-Ju Qian, Tong-Tong Liang et al.· 0 citations
Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training reduces exposure bias, yet finite rollouts leave long-range drift unresolved. We observe that the noise level of the history key-value (K/V) representations trades visua...
Weiqiang Wang, Zhuo-Kun Chen, Yu-Sheng Dai et al.· 0 citations