Improving Policy Learning via Language-Guided State Representation in World Models
Abstract
World models have emerged as powerful technology for facilitating policy learning in robotics by providing predictive representations of environmental dynamics. A crucial component of such models is the internal state representation, which serves as a bridge between observation, decision-making, and future state prediction. However, in multi-task settings, the presence of task-irrelevant information within the representation, especially under limited information capacity, can significantly impair the effectiveness of policy learning. To address this challenge, we propose Language-Guided State Representation (LGSR), a novel method that leverages language-based task descriptions as priors to guide the encoding of state representations within the world model. By incorporating language instructions, LGSR promotes the extraction of task-relevant features while suppressing irrelevant information, thereby enhancing the alignment between representation learning and task objectives. We evaluate LGSR on the CALVIN benchmark, a suite of language-conditioned robotic manipulation tasks. Experimental results demonstrate that incorporating language priors into the state representation enhances the extraction of task-relevant information, thereby significantly improving the efficiency and success rate of downstream policy learning.