These findings motivate Action-Relevant Predictive States (ARPS), a compact predictive interface between the video and action experts that achieves 99.2% success on LIBERO and transfers to LIBERO-Plus without adaptation, reaching 87.3% and exceeding Fast-WAM by 39.2 percentage points.
Qiwen Gu, Jifan Li, Bing-Jie Gao et al.· 0 citations
High similarity between first-visit and return frames does not necessarily show that a video world model remembered the scene; the intervening rollout may simply have changed very little. This ambiguity makes absolute revisit scores sensitive to rendering stability, repetitive content, and failed motion. We introduce \...
Qiwen Gu, Bingjie Gao, Rui Chen et al.· 0 citations
Inspired by render-based compression, this work renders textual chains of thought into images, extract visual features, and construct a discrete latent vocabulary via clustering-based fine-tuning, and concludes that discrete latent tokens provide a controllable and interpretable basis for efficient latent reasoning.
Shuochen Chang, Qingyang Liu, Shaobo Wang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.