Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop ha...
Shi-Feng Bao, Fan-Ding Huang, Yi-Han Lin et al.· 0 citations
DA-Nav is proposed, a Direction-Aware vision-language Navigation framework that reformulates navigation as a discrete spatial grounding problem on the egocentric 2D image plane, outperforming existing State-of-The-Art (SoTA) methods while maintaining a substantially stronger recovery capability.
Ye Yuan, Kehan Chen, Xinqiang Yu et al.· arXiv.org· 0 citations
JEPA-WAM, a latent WAM built in a pretrained V-JEPA space, which couples latent transition prediction with continuous action generation through a shared predictor, predicts a spatially structured joint current-future target that captures task-shared visual temporal structure between current and future observations, whi...
Yi-Han Lin, Jia-Wei He, Shi-Feng Bao et al.· 10 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.