Instance Image-Goal Navigation (IIN) asks an agent to locate the specific object instance shown in a goal image. Existing 3D Gaussian Splatting (3DGS) based methods rely on pose-centric search—sampling many viewpoints, rendering them, and comparing against the goal—which is inefficient in continuous 6-DoF space. We ins...
Yijie Deng, Shuaihang Yuan, Geeta Chandra Raju Bethala et al.· IEEE Robotics and Automation...· 4 citations· ⚡1
Embodied visual tracking requires a robot to choose actions that keep a moving target observable at a suitable distance, and to recover it after occlusion, out-of-view drift, or distractor crossings. We cast the task as planning over future target evidence with an action-conditioned world model. In logged tracking data...
Jun-Yi Hu, Shuaihang Yuan, Jia-Zhao Liang et al.· 0 citations
SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams, provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evalu...
Jun-Yi Hu, Zhe-Wen He, Hao Huang et al.· 0 citations
VTaMo is presented, a framework that introduces explicit multi-granularity alignment at three levels: local alignment via entropy-regularized optimal transport with a learnable null token for fine-grained frame-to-token correspondences; global alignment via a learnable orthogonal transformation that calibrates embeddin...