This work introduces additional 3D spatial prior information into both the training and inference stages of video diffusion models to enhance the spatial structural consistency of generated videos and introduces an energy-function guidance strategy (Warp-Guidance) driven by warping priors during denoising.
Beyond semantic content, camera parameters play a pivotal role in dictating the geometric perspective and appearance of any given image. While recent image editing models excel at semantic and stylistic manipulation, they struggle with explicit camera parameter control. When handling large perspective shifts, instructi...
Xin Shen, Chengyou Jia, Ke Xing et al.· 1 citation
This work proposes an adapter-based framework that incorporates event-derived cues into a pre-trained image-to-video diffusion model with minimal architectural changes and consistently outperforms existing state-of-the-art approaches.
Gui-Xu Lin, Yu-Yang Yu, Xiang Ji et al.· 0 citations
Diffusion models with prompt and reference image-guided editing have seen rapid progress, yet they remain too coarse for pixel-level control. One promising direction is to incorporate a soft mask that specifies spatially varying edit strengths but training such fine-grained control demands expensive pixel-wise annotati...
Digital zoom transitions between dual cameras often exhibit conspicuous discontinuities in geometric structure and chromatic consistency, degrading the user experience. While recent dual-camera smooth zoom (DCSZ) methods attempt to mitigate this by fine-tuning frame interpolation (FI) models on DCSZ data, they struggle...
Jia-Yi Zhang, Ren-Rong Wu, Yu-Kang Ding et al.· 0 citations
This work repurposes pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task, and inherits naturally structured knowledge and richer priors from the video model, enabling more data efficient and effective learning of...
Haosen Yang, Ji-Fei Song, Zhensong Zhang et al.· 1 citation
Observation Operator Diffusion is proposed, a unified framework that aligns both the supervision trajectory and feature refinement with the intrinsic recovery order of image structures and introduces GL-CoDA, a decoder that injects scale-specific Gaussian-Lanczos observations across decoding stages for coarse-to-fine f...
Shaojie Guo, Li-Chen Ma, Haoyang Tong et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.