This work shows that reasoning for control policies can emerge by organizing information in an autoregressive latent space reminiscent of a memory palace, where retrieval is iterative and adaptive.
Abstract
Human decision-making is highly flexible -- some actions are taken immediately; others require longer deliberation. Language models have exhibited a similar capacity for adaptive"reasoning."However, transferring this capability to continuous control policies has been challenging, as directly reasoning in language space may lack the granularity for spatial understanding and precise motions. In this work, we show that reasoning for control policies can emerge by organizing information in an autoregressive latent space reminiscent of a memory palace, where retrieval is iterative and adaptive. Our method, Latent Memory Palace (LMP), formulates reasoning as variational inference with an autoregressive latent distribution. We derive a latent-space reinforcement learning technique to tractably optimize its variational lower bound. The resulting policy, LMP-$\pi$, achieves strong empirical performance in simulation and real-world domains while exhibiting interpretable, adaptive allocation of test-time compute. We further show that the same framework yields a variable-length action tokenizer, LMP-$\texttt{tok}$, which significantly improves the performance of downstream autoregressive policies. Together, these results present a new perspective on latent reasoning for control through the lens of variational inference.
InfRL (Inference-time Reinforcement Learning) offers a practical and domain-agnostic approach to harness reinforcement learning during inference, bridging the gap between static prompting and computationally intensive parameter-level fine-tuning.
Sikun Guo, Amir Hassan Shariatmadari, Jiuqi Wang et al.· Proceedings of the 32nd ACM...· 0 citations
It is shown that score-conditioned In-Context Learning (ICL) admits a structural correspondence to policy gradient optimization, and an exact upper bound on the distribution shift induced by a bounded attention update is derived, yielding a trust-region-like analogy to KL-constrained policy optimization.
Behavior policies are often formulated as continuous generative models, whose iterative denoising processes are expressive but difficult to interpret and prone to producing implausible actions. We propose the Large Discrete Policy (LDiP), a fully discrete behavior modeling framework that selects actions from a large vo...
Zhen-Xin Li, Nadine Chang, Xing-Long Sun et al.· 1 citation
This work proposes "raise one and infer three" diffusion policy (ROITDP), a novel approach that introduces two complementary mechanisms, including a reasoning mechanism built upon the Chain-of-Skill Noise Watermark, which enables temporally coherent multi-step reasoning throughout the diffusion process under distributi...
Yi-Hang Zhu, Yuxuan Wang, Tong Li et al.· Proceedings of the Thirty-Fi...· 0 citations
Readout Feedback (RoFB), a test-time intervention that converts intermediate predictions into token-wise pairwise coupling forces injected into the latent dynamics, is introduced, suggesting that closed-loop steering of latent dynamics can serve as a complementary inference-time control mechanism for recurrent reasonin...
Shunsuke Kamiya, Masanori Koyama, Seongcheol Jeong et al.· 0 citations
The World Critic Model is proposed, built on a lightweight LeJEPA architecture; WCM jointly predicts future latent state and estimates values, such that the critic's representation is explicitly trained to capture temporal dynamics rather than merely regress scalar returns.
Senyu Fei, Xiaopeng Yu, Siyin Wang et al.· arXiv.org· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.