Skip to content
Conference

Robust 3D-Aware Video Object Tracking for Mobile Robots Based on 2D Vision Foundation Models

Jul 2026 · 2026 23rd International Conference on Ubiquitous Robots (UR) · pp. 534-539 · 0 citations · 23 references
Computer Science

Abstract

Egocentric cameras are widely used in robotic navigation and manipulation, yet conventional 2D Video Object Tracking (VOT) methods suffer from severe performance degradation under rapid viewpoint changes and frequent frame-out events. Because most existing trackers rely solely on 2D appearance cues, they often fail to recover object identities once targets temporarily disappear. We propose R3DVOT, a 3D-aware tracking framework that augments 2D vision foundation models with spatial geometric reasoning. Its core component, the Position-Aware Memory Selection (PAMS), lifts mask candidates into a canonical 3D world coordinate system and maintains a persistent world-frame state for position-consistent hypothesis selection. By curating the memory bank with spatially consistent anchors, R3DVOT improves robustness to occlusion and frame-out events. On the VOT benchmark, R3DVOT achieves an AUC improvement of 5.1% over SAM 2 and 1.4% over SAMURAI. Furthermore, on the VOS benchmark, R3DVOT increases the J&F score by 4.6% compared to SAM 2 and 9.0% compared to SAMURAI. These results highlight the effectiveness of 3D spatial continuity in enhancing tracking, segmentation, and long-term identity consistency in robotic perception.

View source

Similar papers

Preprint Aug 2026

EgoTrack3D: A Modular Framework for Egocentric 3D Object Tracking

Understanding 3D scenes from egocentric video is fundamental for robotics and autonomous navigation, yet rapid viewpoint changes and partial occlusions make building structured representations challenging. Existing 3D tracking and scene graph construction methods primarily address explicit interactions or assume static...

Jan Kulik, Bjarni Dagur Thor Karason, Yung-Hsu Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TFTrack: A Template-Free Framework for Efficient 3D Point Cloud Tracking

LiDAR-based 3D Single Object Tracking (3D SOT) is critical for robotic perception and navigation and aims to localize dynamic objects across frames in sparse point clouds. Existing methods, rooted in the Siamese tracking paradigm from 2D vision, rely on costly dual-input designs and excessive motion modeling guided by...

Zhao-Feng Hu, Si-Fan Zhou, Jia-Hao Nie et al. · 2 citations
Preprint Sep 2026

Re-engineering SORT-based algorithms for low-cost small object tracking from omnidirectional footage

Multi-object tracking (MOT) has advanced rapidly in urban surveillance and autonomous driving, yet many trackers rely on ReID- and transformer-based appearance encoders and are designed for standard FoV cameras. These assumptions break down for low-cost omnidirectional deployments, where equirectangular projection intr...

Xin Shu, Meegan Gower, Y. Buckley et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Tracking the Unseen: An Occlusion-Robust Framework for Target Tracking Under Full and Long-Term Occlusion

Real-time multi-object tracking systems remain highly vulnerable to full and long-term occlusion, where targets temporarily or completely disappear from the camera's field of view. Conventional trackers may terminate trajectories prematurely, resulting in identity loss and reduced situational awareness in applications...

Mais.M Mohammed, Sharifa Mohammed, Hanan Awadh et al. · 0 citations
Conference Aug 2026

Monocular Visual-Inertial Odometry via Explicit 3D Motion Modeling

Visual-inertial odometry (VIO) estimates motion by fusing camera and inertial data. In learning-based monocular VIO, depth, scale, and motion are tightly coupled, while IMU cues are difficult to impose as geometric constraints on implicit 2D features. This paper presents a pose estimation method based on explicit 3D mo...

Qian Li, Fan Bai · 0 citations
Preprint Sep 2026

VidAct: Learning Manipulation from In-the-Wild Videos with Object-Centric 3D Awareness

Video demonstrations offer a scalable alternative to costly robot data for learning manipulation, yet existing reconstruction-based approaches often rely on constrained camera viewpoints or human-to-robot retargeting, while the reconstructed trajectories are difficult to adapt to new objects configurations without dist...

Hang Li, Ming-Xin Zhang, Zi-Han Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.