Vision-Language-Action (VLA) or World Action (WAM) models have recently demonstrated remarkable performance in robotic manipulation. On LIBERO, SOTA method have achieved nearly 100\% success rates, seemingly suggesting that the models are ready for deployment in real world. However, near perfect performance on existing...
Lin Liu, Zhi-Cheng Bao, Lu Zhang et al.· 4 citations
Prefix-Steered Recurrent Memory (PREM), a memory-token-free framework for frozen vision-language models (VLMs) that separates video ingestion from query answering, consistently outperforms frozen baselines at every evaluated visual budget.
Si-Ru Zhong, Qiong-Yan Wang, Xiao-Hui Lv et al.· 0 citations
PILOT (Physical Inference for Latent Optimized Trajectories) bridges the gap between high-level physical condition evolution and low-level action trajectory generation within the Action Model, creating a structural bottleneck while weakening the predictive capability of world evolution modeling for action generation.
Xiangkai Ma, Yue Ma, Junjie Wang et al.· 0 citations
This work formalizes Observable Simulator Contract, a minimal contract that any action-conditioned physical simulator should satisfy: supplied actions must induce corresponding agent motion, and environment responses must be grounded in that realized motion.
P. Co, Sichen Hu, Chun-Xuan Jiao et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.