The Geometry Grounded Tracking Anything Model is proposed, a unified framework for promptable instance tracking in 3D using only unordered RGB images or videos and delivers strong cross-view consistency, promptable instance spatial tracking, video object segmentation and spatial reconstruction, establishing a foundation for interactive, geometry-grounded spatial reasoning.
Abstract
Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current video segmentation models depend on explicit object appearance memory banks for instance tracking, yet they remain vulnerable to large viewpoint changes and long-term occlusions. Leveraging the spatial consistency afforded by modern feed-forward 3D reconstruction models, we propose the Geometry Grounded Tracking Anything Model (G$^2$TAM), a unified framework for promptable instance tracking in 3D using only unordered RGB images or videos. G$^2$TAM employs spatially aligned geometric representations as implicit memory, ensuring stable instance identity and localization across frames and views. At its core is a cross-modal spatial encoder that integrates visual and textual prompts into a shared geometric space, enabling end-to-end spatial reconstruction and instance-consistent mask prediction. To support training and evaluation, we construct InsTrack, a large-scale dataset with a dedicated validation split for benchmarking. Extensive experiments show that G$^2$TAM delivers strong cross-view consistency, promptable instance spatial tracking, video object segmentation and spatial reconstruction, establishing a foundation for interactive, geometry-grounded spatial reasoning.
Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms existing streaming baselines while maintaining scalable online inference for long dynamic sequences.
Qwen-3D is introduced, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes and incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene repre...
Lucy Lin, Ayush Jain, Yifan Liu et al.· 3 citations
GaussianWAM is proposed, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field and improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation.
Zi-Jian Zhang, Yu-Qing Jiang, Wei-Tao Zhou et al.· 1 citation
This work introduces PLANET, an end-to-end multi-object tracker designed to move beyond the image plane, and lifts existing 2D tracking datasets into 3D by embedding reconstructed 3D scene geometry into the features and positional encodings used during query formation.
Orcun Cetintas, Guillem Brasó, Tim Meinhardt et al.· 0 citations
Object-centric video models represent scenes with slots, yet exposed geometry can vary in meaning with appearance. In Invariant Slot Attention (ISA), explicit position and scale can disagree with the decoded center and extent; edits can yield unexpected motion or resizing, and replacing appearance can shift geometry. G...
Hao Huang, Zhe-Kai Wang, Xiang Liu et al.· 0 citations
PhysMLLMs is a training-stage prior injection architecture that injects physics-inspired spatial continuity priors into Video MLLMs, demonstrating that the injected spatial prior improves video consistency without compromising image-level grounding or general multimodal capability.
Siyao Yan, Bo Han, Jisheng Dang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.