Jul 2026
τ: Learning Touch-Augmented Vision-Language-Action Models from Future Visual Supervision
Tau is presented, a touch-augmented VLA framework that learns an action-conditioned spatiotemporal tactile representation from future visual supervision inspired by the Joint-Embedding Predictive Architecture (JEPA), and fuses it with vision-language features for action generation.
Ning Cheng, Jinan Xu, Wanlin Li et al.
· arXiv.org · 1 citation