Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· pp. 8341-8342· 0 citations· 11 references
TL;DR
This thesis is developing a context-adaptive Mamba-based world model that infers K context vectors from short pixel-trajectory windows, and shows that its suboptimality is bounded by √K times the context-inference error along a Lipschitz chain through FiLM, the SSM, and the decoder.
Abstract
A vision-based robot can be faced with changing friction, gravity, or motor behavior, whose effects are observable only from pixels mid-episode. Visual model-based reinforcement learning struggles in such settings because most world models assume the latent factors driving these changes stay fixed during interaction, or fold them into a single context embedding. My thesis is developing a context-adaptive Mamba-based world model that infers K context vectors from short pixel-trajectory windows. Each vector steers a disjoint slice of the Mamba backbone through block-diagonal feature-wise linear modulation (FiLM), so distinct dynamics signals travel along parallel pathways rather than collapse into a single channel. The policy trains on imagined rollouts conditioned on the inferred context, and I show that its suboptimality is bounded by √K times the context-inference error along a Lipschitz chain through FiLM, the SSM, and the decoder. Evaluation spans within-episode shifts on DeepMind Control (Walker, Quadruped) and mode/difficulty switching on Atari, against an inferred-context baseline and a privileged-context oracle.
Although highly effective in vision and language domains, applying in-context learning to robotics remains challenging. Existing autoregressive in-context imitation methods discretize continuous actions and exacerbate the accumulation of early prediction errors through next-token prediction, limiting their generalizati...
Jian Ding, Xian-Jie Dai, Roei Herzig et al.· 0 citations
World Action Models (WAMs) acquire behavioral priors by modeling future scene evolution, but predicting detailed futures in pixel or latent space incurs substantial cost. Recent evidence that co-training gains persist without test-time generation raises a question: what must a WAM learn to improve control? We introduce...
Ze Feng, Yi-Xu Feng, Ling-Yu Xiao et al.· 0 citations
World-model planners typically scale outward by rolling farther, sampling more trajectories, or optimizing longer, while assigning the same computation to every imagined transition. We show that making every transition uniformly deeper wastes computation and can degrade planning because useful refinement is concentrate...
Zi-Jian Jin, Yun-Bei Zhang, Yuan-Zhe Liu et al.· 0 citations
World-action models jointly learn robot policies and predict future observations, making the representation space an interface between control and prediction. We study the design of this space through controlled comparisons, finding that neither reconstruction fidelity nor pre-trained perceptual features alone ensure e...
Hao-Yi Jiang, Liu Liu, Xin-Jiang Wang et al.· 0 citations
Vision-language-action models often predict actions from only the current observation, which can leave tasks involving object occlusion or visually identical objects ambiguous without episode history. The usual countermeasure, widening the observation window, turns the horizon into a hyperparameter and lets per-step co...
WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token is introduced and its data-scaling and temporal-context behavior under the tested recipes are characterized.
Chunkai Yang, An-Dong Yang, Di Huang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.