Skip to content
Preprint

MV2: Multi-View Multi-Vehicle Driving Dataset for Novel View Synthesis

Aug 2026 · 0 citations · 46 references
Computer Science

TL;DR

benchmarking recent NVS and camera pose estimation methods shows that NVS performance degrades with increasing viewpoint disparity, and that feed-forward pose estimators notably lag behind optimization-based approaches, highlighting MV2 as a rigorous testbed for NVS in driving.

Abstract

Differentiable rendering has advanced novel view synthesis (NVS), yet applying it to real-world driving remains difficult due to sparse capture viewpoints, dynamic objects, and limited multi-trajectory data. We introduce the Multi-View Multi-Vehicle (MV2) dataset and benchmark for evaluating NVS models under large viewpoint changes in dynamic urban scenes. MV2 features synchronized captures from a car, scooter, and drone, each following distinct yet synchronized trajectories. Training NVS methods on one vehicle's camera stream and testing on another enables evaluation under substantially larger viewpoint variations than existing single-trajectory datasets. All sequences are registered via Structure-from-Motion and camera poses verified using manual pixel-level correspondence annotations, yielding 50 high-quality scenes with 12000 images. Benchmarking recent NVS and camera pose estimation methods shows that NVS performance degrades with increasing viewpoint disparity, and that feed-forward pose estimators notably lag behind optimization-based approaches, highlighting MV2 as a rigorous testbed for NVS in driving. The dataset, benchmark protocol, and project resources are available at https://mv2-dataset.github.io/.

View source

Similar papers

Preprint Sep 2026

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

This work introduces TAPVid-MV (Tracking Any Point in Video across Multiple Views), the first benchmark for multi-view 3D point tracking, and identifies geometry recovery as a major bottleneck for accurate 3D point tracking.

Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman et al. · 1 citation
Preprint Sep 2026

M3GD: Multi-Modal Multi-View Geometric Diffusion for Camera--LiDAR Novel View Synthesis

Robotic novel view synthesis (NVS) must recover both visual appearance and metric 3D structure, yet most generative NVS methods rely only on images, overlooking LiDAR, a complementary sensor common on robotic platforms. We present M3GD, a Camera--LiDAR multimodal representation for generative NVS that composes independ...

Yang Zhou, Jiuhong Xiao, Shi-Zhao Ye et al. · 0 citations
Preprint Aug 2026

PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction

PIVOT (Pose, Intrinsics and Viewpoint Oriented Testbed), a multi-trajectory dataset, processing pipeline, and evaluation framework for independently studying novel-view synthesis methods, is introduced and a directed pose-space Chamfer distance is introduced to quantify how well training poses cover an evaluation traje...

M. Raymond · 0 citations
Preprint Sep 2026

VDGS: Visibility-Driven Large-Scale 3D Gaussian Splatting for Aerial Scene Reconstruction

Large-scale scene reconstruction is a critical foundational technology in robotic autonomous systems such as 3D mapping and autonomous driving. In recent years, 3D Gaussian Splatting (3DGS) has demonstrated remarkable advantages in both visual quality and computational efficiency, making it a promising representation f...

Hao-Lin Yu, Jia-Dong Tang, Yi-Xian Wang et al. · 0 citations
Preprint Aug 2026

R2S-EGO: Dual-Proxy Refinement for Sparse-Capture Real-to-Sim

Real-to-sim (R2S) depends on scene representations that render observations along robot ego trajectories, yet dense multi-view capture limits per-environment real-image capture-count efficiency, and sparse human capture can leave behavior-scoped robot views under-supported. Camera-controlled synthesis can fill missing...

Shuai Fang, Xin Deng, Yuchen Kang et al. · 1 citation
Aug 2026

Dynamic View Synthesis from Monocular Videos via Motion-aware Gaussian Splatting.

This paper proposes a semantics-guided scene decoupling module that separates Gaussian primitives into static and dynamic components based on motion vectors, and introduces a motion-aware densification module for motion compensation, which alleviates the incomplete rendering of dynamic objects caused by insufficient sp...

Chulin Zhao, Xue Wang, Guo-Qing Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.