Skip to content

G2TAM: Geometry Grounded Track Anything Model

Jul 2026 · arXiv.org · Vol abs/2607.03789 · 0 citations · 62 references
Computer Science

TL;DR

The Geometry Grounded Tracking Anything Model is proposed, a unified framework for promptable instance tracking in 3D using only unordered RGB images or videos and delivers strong cross-view consistency, promptable instance spatial tracking, video object segmentation and spatial reconstruction, establishing a foundation for interactive, geometry-grounded spatial reasoning.

Abstract

Human spatial understanding arises from jointly perceiving geometry and semantics, enabling consistent object identification and localization across viewpoints and time. Current video segmentation models depend on explicit object appearance memory banks for instance tracking, yet they remain vulnerable to large viewpoint changes and long-term occlusions. Leveraging the spatial consistency afforded by modern feed-forward 3D reconstruction models, we propose the Geometry Grounded Tracking Anything Model (G$^2$TAM), a unified framework for promptable instance tracking in 3D using only unordered RGB images or videos. G$^2$TAM employs spatially aligned geometric representations as implicit memory, ensuring stable instance identity and localization across frames and views. At its core is a cross-modal spatial encoder that integrates visual and textual prompts into a shared geometric space, enabling end-to-end spatial reconstruction and instance-consistent mask prediction. To support training and evaluation, we construct InsTrack, a large-scale dataset with a dedicated validation split for benchmarking. Extensive experiments show that G$^2$TAM delivers strong cross-view consistency, promptable instance spatial tracking, video object segmentation and spatial reconstruction, establishing a foundation for interactive, geometry-grounded spatial reasoning.

View source

Similar papers

Jul 2026

IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

Experiments on 3D reconstruction, pose estimation, instance spatial tracking, and open-vocabulary segmentation demonstrate that IGGT4D outperforms existing streaming baselines while maintaining scalable online inference for long dynamic sequences.

Zheng-Yu Zou, Hao Li, Kuixuan Jiao et al. · 1 citation
Preprint Aug 2026

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding

Qwen-3D is introduced, a geometry-aware LMM that compresses visual information within the Qwen backbone using multi-view geometric cues, enabling efficient long-horizon visual reasoning over static scenes and incorporates a query-based segmentation decoder that grounds language directly in the underlying 3D scene repre...

Lucy Lin, Ayush Jain, Yifan Liu et al. · 3 citations
Preprint Aug 2026

GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models

GaussianWAM is proposed, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field and improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation.

Zi-Jian Zhang, Yu-Qing Jiang, Wei-Tao Zhou et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Beyond the Image Plane: World-Grounded Queries for Multi-Object Tracking

This work introduces PLANET, an end-to-end multi-object tracker designed to move beyond the image plane, and lifts existing 2D tracking datasets into 3D by embedding reconstructed 3D scene geometry into the features and positional encodings used during query formation.

Orcun Cetintas, Guillem Brasó, Tim Meinhardt et al. · 0 citations
Preprint Sep 2026

GeoCo-SAVi: Geometry-Consistent Slot Attention for Explicitly Editable Object Representations

Object-centric video models represent scenes with slots, yet exposed geometry can vary in meaning with appearance. In Invariant Slot Attention (ISA), explicit position and scale can disagree with the decoded center and extent; edits can yield unexpected motion or resizing, and replacing appearance can shift geometry. G...

Hao Huang, Zhe-Kai Wang, Xiang Liu et al. · 0 citations
#small language model Preprint Aug 2026

PhysMLLMs: Spatial Priors for Unified Referring Segmentation and Grounded Reasoning of Images and Videos

PhysMLLMs is a training-stage prior injection architecture that injects physics-inspired spatial continuity priors into Video MLLMs, demonstrating that the injected spatial prior improves video consistency without compromising image-level grounding or general multimodal capability.

Siyao Yan, Bo Han, Jisheng Dang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.