Skip to content

Author

Zi-Tai Huang

We have 3 of 8 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

CSWAM: Better Causal Semantic Representations for Out-of-Distribution Generalization in World Action Models

FastWAM-style world action models enable efficient action-only inference, but generalize poorly under visual distribution shifts. Their reconstruction-oriented representations emphasize appearance-specific details, limiting generalization to unseen scenes and objects. Without observation history, the model also lacks temporal evidence for robustly identifying task-relevant state changes and motion in unfamiliar visual conditions. To address these limitations, we present the Causal Semantic World Action Model (CSWAM), which augments FastWAM with a causal semantic expert built on V-JEPA 2.1. V-JEPA provides temporally grounded representations of semantic state changes and motion with less dependence on appearance-specific details. The expert learns their future evolution from a sparse history of current and past observations and shares the history-derived context with both the video and action streams through causal attention. At inference, CSWAM conditions action denoising on the current video state and observed semantic history, retaining efficient action-only inference. We conduct simulation and real-robot experiments to evaluate generalization under distribution shifts. With embodied pretraining, CSWAM raises Randomized success on RoboTwin 2.0 Clean-to-Randomized transfer from 10.16% to 45.18%, a gain of 35.02 percentage points over FastWAM. Across two real-robot tasks and three OOD difficulty levels, CSWAM improves average success over FastWAM by 42.5 percentage points, from 27.5% to 70.0%.

Tian-Bin Liu, Jian Zhu, Taiyi Su et al. · 0 citations
Jul 2026

DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation

DSWAM is introduced, a Dual-System World Action Foundation Model for fine-grained robot manipulation and built and evaluated under the DeMaVLA real-world deformable manipulation setting with matched robot platform, pretraining data, post-training data, and evaluation criteria.

Jian Zhu, Jianjun Zhang, Taiyi Su et al. · 2 citations
Preprint Jul 2026

Learning 4D Geometric Priors for Inference-Efficient World Action Models

MECo-WAM is proposed, a Multi-Expert Co-Training World Action Model that injects action-relevant 4D geometric priors into video-action representations while preserving the original lightweight inference graph.

Jianjun Zhang, Jian Zhu, Taiyi Su et al. · 5 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.