Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the vis...
Yuan Xu, Yi-Xiang Chen, Qi-Sen Ma et al.· 0 citations
BridgeVLA++ is developed by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history that can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities.
A systematic shortcut audit of EmoPrefer using content-blind probes shows that the current scores can be reached without verifying either description against the video, and recommends source-balanced pairing, strict length control, counter-stereotypical sliced reporting, and multi-annotator consensus for future cross-g...
XEWorld is introduced, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying ph...
Yixiang Chen, Jiabing Yang, Yuan Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.