Jul 2026· Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Talks· pp. 1-3· 0 citations· 2 references
TL;DR
A LiDAR-Constrained Deep Visual Odometry system, a robust tracking architecture designed to solve these "impossible" shots by fusing pre-existing LiDAR geometry with modern Deep Learning, and introduces a "Leapfrogging" architecture that automatically detects and corrects temporal drift by re-anchoring to the geometry from trusted keyframes.
Abstract
Matchmoving is the bedrock of visual effects, yet it remains a fragile bottleneck when footage contains heavy motion blur, low texture, or dynamic occlusion. While physical on-set camera tracking (e.g., encoded cranes) exists, it is often impractical for handheld interior shots and prone to mechanical slippage, leaving post-production software to solve the gap. This talk presents a LiDAR-Constrained Deep Visual Odometry system, a robust tracking architecture designed to solve these "impossible" shots by fusing pre-existing LiDAR geometry with modern Deep Learning. Unlike traditional commercial solvers that hunt for sparse, high-contrast corners, our approach uses Deep Optical Flow (RAFT) to track the entire dense image context, locking the camera directly to the set’s 3D mesh. We introduce a "Leapfrogging" architecture that automatically detects and corrects temporal drift by re-anchoring to the geometry from trusted keyframes. By prioritizing geometric truth over feature quantity, this standalone Python tool reduces days of manual hand-tracking and rotoscoping to minutes of automated computation, achieving high median precision on sequences where standard algorithms fail entirely.
This work introduces a system that combines the strong sequential constraints of SLAM with the flexibility and global optimization of offline SfM, enabling the metric reconstruction of arbitrary, long, uncalibrated videos.
Zador Pataki, Paul-Edouard Sarlin, Marc Pollefeys· arXiv.org· 1 citation· ⚡1
A unified calibration-free 3D MCPT framework that infers geometric structure directly from visual data using deep foundation models is proposed, establishing a strong baseline for purely vision-based 3D tracking.
This work introduces FastEventDGS, a novel Deformable Gaussian Splatting-based framework that leverages a single event camera for high-fidelity 4D reconstruction in dynamic scenes and proposes a local patch event motion loss to constrain object motion, effectively mitigating over-fitting.
Zijia Dai, Nico Messikommer, Rong Zou et al.· 0 citations
KP-SLAM is proposed, which predicts dense optical flow and paired pointmap priors from a shared representation and incorporates them into the same BA backend and introduces a Depth-Scale-Pose-to-Pointmap (DSPP) objective that relates optimized inverse depth, edge-wise relative scale, and camera pose to paired pointmap...
Song Gao, Xinyu Huang, Zheng Huang et al.· Symmetry· 0 citations
A complete, reproducible RGB-only pipeline and ablation for geometry-first and estimated-depth pseudo-LiDAR is used as a controlled test of one hypothesis: that cross-view geometric consistency, not monocular depth accuracy, governs performance under Sim2Real.
Abdullah Naeem, Anav Katwal, Ayon Dey et al.· 0 citations
Multi-object tracking (MOT) has advanced rapidly in urban surveillance and autonomous driving, yet many trackers rely on ReID- and transformer-based appearance encoders and are designed for standard FoV cameras. These assumptions break down for low-cost omnidirectional deployments, where equirectangular projection intr...
Xin Shu, Meegan Gower, Y. Buckley et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.