World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it...
Wen-Bo Chen, Tian-Fu Li, Hao-Xuan Xu et al.· 0 citations
Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often requires expensive i...
Tian-Fu Li, Hao-Xuan Xu, Wen-Bo Chen et al.· 0 citations
This work proposes an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation, and introduces a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly o...
Lin-Wei Zheng, Dao-Jie Peng, Bing-Tao Wang et al.· 0 citations
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction d...
Jun-Feng Li, Junjie He, Zhi-De Zhong et al.· 1 citation
4D-WAM is proposed, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment, enabling WAMs to learn trajectory-level spatiotemporal representations.
Lishan Yang, Wen-Xuan Song, Xi Wang et al.· 5 citations· ⚡1
SSMB is presented, a deblur-free, self-supervised keypoint detector for motion-blurred images that requires neither handcrafted detectors nor external pseudo-labels, and introduces the Local Discriminability Enhancement (LDE) module, which restores fine-grained local discriminability after global feature mixing.
Zhenjun Zhao, F. Bellavia, Wen-Ting Wang et al.· 0 citations
MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization.
Ze-Hua Fan, Jun-Jie He, Wen-Xuan Song et al.· 3 citations· ⚡1
DreamTrajectory is presented, a trajectory-guided framework for language-conditioned mobile manipulation that introduces one component for each limitation of existing Vision-Language-Action policies, and jointly predicts an intention-level end-effector trajectory and a whole-body action chunk in a single action expert.
Zheng Yang, Wen-Jie Zhang, Xiang-Yu Chen et al.· 1 citation
The Robust-WAM is a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream to retain the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics.
Hao-Dong Yan, Jun-Feng Li, Junjie He et al.· 0 citations
PSG-JEPA is proposed, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes.