Skip to content

Author

Shanghang Zhang

We have 5 of 30 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation

A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a c...

Chu-Yao Fu, Xiao-Wei Chi, Yu-Han Rui et al. · 0 citations
Jul 2026

Hierarchical Denoising For Multi-Step Visual Reasoning

This work proposes HDR (Hierarchical Denoising for Visual Reasoning), a unified framework that integrates hierarchical latents into causal video generation for multi-step reasoning and introduces a level-stratified multi-step video reasoning benchmark with out-of-distribution cases.

Ze-Zhong Qian, Xiao-Wei Chi, Chak-Wing Mak et al. · 1 citation
#artificial intelligence Preprint Sep 2026

DeCAL: Towards Physically-Grounded Dexterous Vision-Language-Action Models via Contact-Aware Latent Co-Imagination

DeCAL is presented, a physically-grounded dexterous vision-language-action model that unifies understanding, imagination and action generation for contact-rich dexterous manipulation and introduces Adaptive Visuo-Tactile Fusion that dynamically regulates tactile interactions via a contact-aware gating strategy.

Yan-Kai Fu, Ning Chen, Jun-Kai Zhao et al. · 0 citations
Jul 2026

Data Pyramid for Embodied Manipulation

This work organizes the embodied data ecosystem as a pyramidspanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelit...

Yifan Ye, Yankai Fu, Ya-hui Lv et al. · 4 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.