Scaling robotic foundation models requires diverse training data and reliable evaluation environments. Simulation offers a scalable solution, yet existing generation pipelines remain constrained by predefined assets and skills, a disconnect between scene generation and task generation, and limited support for complex e...
Yi-Kai Qin, Yi-Fei Deng, Ming-Jian Liang et al.· 0 citations
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semant...
Wen-Xuan Song, Jia-Yi Chen, Jing-Bo Wang et al.· 1 citation
4D-WAM is proposed, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment, enabling WAMs to learn trajectory-level spatiotemporal representations.
Lishan Yang, Wen-Xuan Song, Xi Wang et al.· 5 citations· ⚡1
PSG-JEPA is proposed, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes.