Skip to content

Stabilizing Camera-Controlled Novel View Synthesis at Inference Time

Prajwal Singh Arjun Badola Seema Kumari Hajime Nagahara Shanmuganathan Raman
Sep 2026 · 0 citations · 65 references
Computer Science

TL;DR

CamTrol++ improves temporal and geometric consistency, downstream 3D reconstruction quality, and generation efficiency over training-free baselines over RealEstate10K and MegaScene, and improves temporal and geometric consistency over training-free baselines.

Abstract

Training-free, camera-controlled novel view synthesis from a single image using pre-trained video diffusion models often becomes unstable under large camera motion and long generation horizons. Existing approaches commonly combine several inference-time components, making it unclear which design choices are most important for stability. We show that the main source of stability is simple. Decomposing camera motion into small autoregressive steps limits per-step geometric distortion and reduces error accumulation. A controlled camera-step study shows that performance remains stable for small motions and degrades more strongly as the per-step motion approaches $18$-$20^\circ$. We further evaluate geometry-constrained spatial attention and low-frequency appearance anchoring as supporting refinements, together with an efficient registration-free warping pipeline. Across RealEstate10K and MegaScene, CamTrol++ improves temporal and geometric consistency, downstream 3D reconstruction quality, and generation efficiency over training-free baselines. The method remains effective for 56-frame generation and under substantial controlled depth corruption. These results show that careful control of camera motion at inference time can substantially improve the stability of camera-controlled novel view synthesis without retraining or modifying the diffusion backbone.

View source

Similar papers

Preprint Sep 2026

CameraEditor: Camera-Controlled Image Editing via Video-Prior Sequential Modeling

Beyond semantic content, camera parameters play a pivotal role in dictating the geometric perspective and appearance of any given image. While recent image editing models excel at semantic and stylistic manipulation, they struggle with explicit camera parameter control. When handling large perspective shifts, instructi...

Xin Shen, Chengyou Jia, Ke Xing et al. · 1 citation
Preprint Sep 2026

FFVO: A Feedforward Pose Decoder for Long-Horizon Visual Odometry

Stable and reliable 4D spatial understanding is fundamental for autonomous driving systems. While feedforward reconstruction networks can estimate camera motion and 3D structure in one pass, pose estimation over long videos remains challenged by computational cost, long-context ambiguity, and temporal instability. To a...

Meng-Li Shih, Shih-Yang Su, Yu-Liang Zou et al. · 0 citations
Preprint Aug 2026

ReX-Shot: Single-Image Rephotography via Geometry- and Camera-Grounded Generation

Single-image rephotography aims to synthesize new shots of a scene from a single reference image with specified viewpoints, focal lengths, and photographic effects, which are intrinsically coupled in imaging. Existing methods typically treat these factors separately and struggle under joint control: novel-view synthesi...

Rui-Qi Zhang, Hao Zhu, Wen-Hao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PartiCam: Camera Controlled Video Generation with Reward Guidance

We present PartiCam, a training-free Particle filtering rooted method for improved Camera controlled video generation. Generating videos that follow a precisely specified camera trajectory remains challenging for large video diffusion models. Training-free approaches are backbone-agnostic and avoid the need to construc...

Amine Ouasfi, Runjia Li, Junlin Han et al. · 0 citations
Preprint Sep 2026

Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation

Long-horizon camera-controlled video generation requires recovering previously observed content from an ever-growing visual history. Existing approaches either search historical context implicitly or reconstruct it into persistent 3D memory, facing inefficient memory access or accumulated geometric errors. Our key insi...

Ze-Song Yang, Wei-Kai Chen, Li-Yuan Cui et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.