Skip to content

Author

Haoang Li

We have 10 of 39 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it...

Wen-Bo Chen, Tian-Fu Li, Hao-Xuan Xu et al. · 0 citations
Preprint Sep 2026

SLIP-VLA: Single-Step Latent Imagination for Policy Learning in Vision-Language-Action Models

Vision-Language-Action models are increasingly effective for robotic manipulation, yet most predict actions directly from current observations without explicitly modeling future scene evolution. Recent methods introduce future prediction to improve action generation, but dense future modeling often requires expensive i...

Tian-Fu Li, Hao-Xuan Xu, Wen-Bo Chen et al. · 0 citations
Preprint Sep 2026

A Topological Representation with Object-Path Graphs for Open-Vocabulary Instance Navigation

This work proposes an object--path graph that unifies open-vocabulary semantic reasoning with topological navigation, and introduces a navigation strategy that combines global path planning with local inter-node execution through lightweight node localization and semantic visual servoing, enabling navigation directly o...

Lin-Wei Zheng, Dao-Jie Peng, Bing-Tao Wang et al. · 0 citations
Preprint Aug 2026

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction d...

Jun-Feng Li, Junjie He, Zhi-De Zhong et al. · 1 citation
Preprint Aug 2026

4D-WAM: Infusing Spatiotemporal Awareness into World Action Models through Trajectory Fields

4D-WAM is proposed, a model-agnostic training strategy that injects spatiotemporal knowledge from 3D trajectory fields into WAMs through representation alignment, enabling WAMs to learn trajectory-level spatiotemporal representations.

Lishan Yang, Wen-Xuan Song, Xi Wang et al. · 5 citations · ⚡1
Preprint Aug 2026

SSMB: Self-Supervised Local Feature Detection under Motion Blur

SSMB is presented, a deblur-free, self-supervised keypoint detector for motion-blurred images that requires neither handcrafted detectors nor external pseudo-labels, and introduces the Local Discriminability Enhancement (LDE) module, which restores fine-grained local discriminability after global feature mixing.

Zhenjun Zhao, F. Bellavia, Wen-Ting Wang et al. · 0 citations
Preprint Aug 2026

DreamTrajectory: Trajectory-Guided Action Generation with World Model Alignment for Mobile Manipulation

DreamTrajectory is presented, a trajectory-guided framework for language-conditioned mobile manipulation that introduces one component for each limitation of existing Vision-Language-Action policies, and jointly predicts an intention-level end-effector trajectory and a whole-body action chunk in a single action expert.

Zheng Yang, Wen-Jie Zhang, Xiang-Yu Chen et al. · 1 citation
Preprint Aug 2026

Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models

The Robust-WAM is a general post-training method for video-generation-based WAMs that preserves the VAE-based generative path and adds a lightweight semantic foresight alignment objective on the action stream to retain the large-scale VGM pretraining while grounding actions in appearance-invariant dynamics.

Hao-Dong Yan, Jun-Feng Li, Junjie He et al. · 0 citations
Preprint Aug 2026

Is Forward Prediction Enough? Physical State Grounding for JEPA World Models

PSG-JEPA is proposed, a physically grounded JEPA world model that shapes its latent space with two complementary grounding objectives beyond forward prediction: grounding individual latents in robot proprioceptive state, and grounding latent pairs in multi-horizon joint-angle changes.

Hao-Dong Yan, Jia-Guang Zhu, Ming-Ming Jia et al. · 6 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.