Embodied visual tracking requires a robot to choose actions that keep a moving target observable at a suitable distance, and to recover it after occlusion, out-of-view drift, or distractor crossings. We cast the task as planning over future target evidence with an action-conditioned world model. In logged tracking data...
Jun-Yi Hu, Shuaihang Yuan, Jia-Zhao Liang et al.· 0 citations
SignDino, a self-supervised sign-video encoder that moves the DINOv3 student--teacher recipe from the spatial domain of image crops to the temporal domain of tracked sign streams, provides a strong public self-supervised representation and shows competitive or state-of-the-art performance under matched downstream evalu...
Jun-Yi Hu, Zhe-Wen He, Hao Huang et al.· 0 citations
VTaMo is presented, a framework that introduces explicit multi-granularity alignment at three levels: local alignment via entropy-regularized optimal transport with a learnable null token for fine-grained frame-to-token correspondences; global alignment via a learnable orthogonal transformation that calibrates embeddin...