Skip to content

Spatial Prior-guided Video Diffusion Models With Controllable Camera Trajectories

· 0 citations · 50 references

TL;DR

This work introduces additional 3D spatial prior information into both the training and inference stages of video diffusion models to enhance the spatial structural consistency of generated videos and introduces an energy-function guidance strategy (Warp-Guidance) driven by warping priors during denoising.

View source

Similar papers

Preprint Sep 2026

CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling

Beyond semantic content, camera parameters play a pivotal role in dictating the geometric perspective and appearance of any given image. While recent image editing models excel at semantic and stylistic manipulation, they struggle with explicit camera parameter control. When handling large perspective shifts, instructi...

Xin Shen, Chengyou Jia, Ke Xing et al. · 1 citation
Preprint Aug 2026

Bridging Event Streams and DiT: Event-Guided Video Frame Interpolation

This work proposes an adapter-based framework that incorporates event-derived cues into a pre-trained image-to-video diffusion model with minimal architectural changes and consistently outperforms existing state-of-the-art approaches.

Gui-Xu Lin, Yu-Yang Yu, Xiang Ji et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Diffusion Editing with Soft Mask: Pixel Level Redo of Image and Video with Adjustable Strength

Diffusion models with prompt and reference image-guided editing have seen rapid progress, yet they remain too coarse for pixel-level control. One promising direction is to incorporate a soft mask that specifies spatially varying edit strengths but training such fine-grained control demands expensive pixel-wise annotati...

Candi Zheng, Yuan Lan · 0 citations
Preprint Sep 2026

ZoomDiff: A High-Fidelity Diffusion Model for Dual-Camera Smooth Zooming

Digital zoom transitions between dual cameras often exhibit conspicuous discontinuities in geometric structure and chromatic consistency, degrading the user experience. While recent dual-camera smooth zoom (DCSZ) methods attempt to mitigate this by fine-tuning frame interpolation (FI) models on DCSZ data, they struggle...

Jia-Yi Zhang, Ren-Rong Wu, Yu-Kang Ding et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Video Generative Models as Geometry Learner

This work repurposes pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task, and inherits naturally structured knowledge and richer priors from the video model, enabling more data efficient and effective learning of...

Haosen Yang, Ji-Fei Song, Zhensong Zhang et al. · 1 citation
Preprint Aug 2026

Pixel-Space Diffusion via Observation Operators

Observation Operator Diffusion is proposed, a unified framework that aligns both the supervision trajectory and feature refinement with the intrinsic recovery order of image structures and introduces GL-CoDA, a decoder that injects scale-specific Gaussian-Lanczos observations across decoding stages for coarse-to-fine f...

Shaojie Guo, Li-Chen Ma, Haoyang Tong et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.