Skip to content
Preprint

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

Sep 2026 · 1 citation · 92 references
Computer Science

TL;DR

This work introduces TAPVid-MV (Tracking Any Point in Video across Multiple Views), the first benchmark for multi-view 3D point tracking, and identifies geometry recovery as a major bottleneck for accurate 3D point tracking.

Abstract

Multi-camera systems are increasingly practical for robotics, AR/VR, and autonomous driving because complementary views reduce depth ambiguity and preserve visibility under occlusion. Existing point-tracking benchmarks, however, focus on a single video or static multi-camera rigs. None test long-term 3D point tracking across several synchronized views under camera motion. We introduce TAPVid-MV (Tracking Any Point in Video across Multiple Views), the first benchmark for this setting. It contains a curated set of 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks across seven subsets spanning indoor and outdoor domains, from robotics and human activity to driving and synthetic procedural scenes. We obtain these trajectories using dataset-specific auxiliary modalities: sensor depth, LiDAR, SLAM and SfM points, human meshes, posed object meshes, and simulation. Every sequence and trajectory is visually verified by human annotators. Across more than 30 baselines, no method comes close to solving the task. Surprisingly, existing multi-view point trackers do not consistently outperform monocular point trackers. By evaluating reconstruction and point tracking on the same datasets, TAPVid-MV helps distinguish errors in recovered geometry from errors in point correspondence. Through this joint analysis, we identify geometry recovery as a major bottleneck for accurate 3D point tracking. Beyond multi-view 3D point tracking, our released annotations support monocular 2D and 3D point tracking, future-trajectory prediction, and 4D reconstruction.

View source

Similar papers

Preprint Aug 2026

MV2: Multi-View Multi-Vehicle Driving Dataset for Novel View Synthesis

benchmarking recent NVS and camera pose estimation methods shows that NVS performance degrades with increasing viewpoint disparity, and that feed-forward pose estimators notably lag behind optimization-based approaches, highlighting MV2 as a rigorous testbed for NVS in driving.

Sanjay Bhargav Dharavath, Hanvitha Saraswathi Mukkamala, F. Khan et al. · 0 citations
Preprint Aug 2026

PIVOT: A Multi-Trajectory Dataset and Testbed for Pose, Intrinsics, and Novel Viewpoint Evaluation in Real-World 3D Reconstruction

PIVOT (Pose, Intrinsics and Viewpoint Oriented Testbed), a multi-trajectory dataset, processing pipeline, and evaluation framework for independently studying novel-view synthesis methods, is introduced and a directed pose-space Chamfer distance is introduced to quantify how well training poses cover an evaluation traje...

M. Raymond · 0 citations
Preprint Aug 2026

Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real

The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) setting; depth is available only for training and validation, so inference is RGB-only. We use two RGB-only routes as a controlled test of one hypothesis: that cross-view geometric con...

Abdullah Naeem, Anav Katwal, Ayon Dey et al. · 0 citations
Preprint Sep 2026

Racing in Volume with Flow Ensembles

Streaming 4D reconstruction has been demonstrated only indoors, on dense camera rigs surrounding subjects that move at human pace. Outdoor 4D reconstruction exists but relies either on cameras mounted on the moving vehicle itself, or on limited-coverage arrays observing quasi-static subjects offline. The case that actu...

Saswat Subhajyoti Mallick, Riu Cherdchusakulchai, Marc Ruiz Olle et al. · 0 citations
Preprint Sep 2026

RoGSW4RLD: Feed-Forward 4D Gaussian Lifting for Robot World Model Rollouts

Action-conditioned video world models predict future robot interactions from multiple cameras, yet their outputs remain disparate video collections rather than a shared metric scene queryable across viewpoints and time. While existing 4D reconstruction methods offer a path to spatialize these predictions, independently...

Jin Hyun Kim, Min Young Kim, Soohwan Song et al. · 0 citations
Preprint Aug 2026

ORBIT++: Benchmarking SfM in the Wild with 360{\deg} Video

A new benchmark for evaluating camera pose estimation is introduced, called ORBIT, to leverage online panoramic 360{\deg} video as a source of data from which to construct challenging clips, while still enabling robust ground-truth trajectory recovery.

S. Sabour, Linyi Jin, Richard Tucker et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.