Vision-Language-Action (VLA) models have become a central paradigm for robot policy learning, which predict actions in three forms: raw action chunks, discrete action tokens, or continuous action latents. However, existing action representations primarily model action trajectories, with limited consideration of the vis...
Yuan Xu, Yi-Xiang Chen, Qi-Sen Ma et al.· 0 citations
BridgeVLA++ is developed by equipping BridgeVLA with a unified spatio-temporal memory architecture that models persistent spatial context and temporal interaction history that can reason over observation histories while preserving BridgeVLA's data efficiency and generalization capabilities.
The Unified Embodied Seeking and Following Benchmark (UESF-Bench) is introduced, a large-scale and diverse benchmark for embodied human seeking and following that requires agents to handle semantic-guided exploration, reliable behavior switching and recovery, and delayed identity grounding.
Kun Yu, Jianhua Yang, Yixiang Chen et al.· arXiv.org· 0 citations
A systematic shortcut audit of EmoPrefer using content-blind probes shows that the current scores can be reached without verifying either description against the video, and recommends source-balanced pairing, strict length control, counter-stereotypical sliced reporting, and multi-annotator consensus for future cross-g...
XEWorld is introduced, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying ph...
Yixiang Chen, Jiabing Yang, Yuan Xu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.