Skip to content

Latent Memory Palace: Reasoning for Control as Autoregressive Variational Inference

Jul 2026 · arXiv.org · Vol abs/2607.08724 · 0 citations · 36 references
Computer Science

TL;DR

This work shows that reasoning for control policies can emerge by organizing information in an autoregressive latent space reminiscent of a memory palace, where retrieval is iterative and adaptive.

Abstract

Human decision-making is highly flexible -- some actions are taken immediately; others require longer deliberation. Language models have exhibited a similar capacity for adaptive"reasoning."However, transferring this capability to continuous control policies has been challenging, as directly reasoning in language space may lack the granularity for spatial understanding and precise motions. In this work, we show that reasoning for control policies can emerge by organizing information in an autoregressive latent space reminiscent of a memory palace, where retrieval is iterative and adaptive. Our method, Latent Memory Palace (LMP), formulates reasoning as variational inference with an autoregressive latent distribution. We derive a latent-space reinforcement learning technique to tractably optimize its variational lower bound. The resulting policy, LMP-$\pi$, achieves strong empirical performance in simulation and real-world domains while exhibiting interpretable, adaptive allocation of test-time compute. We further show that the same framework yields a variable-length action tokenizer, LMP-$\texttt{tok}$, which significantly improves the performance of downstream autoregressive policies. Together, these results present a new perspective on latent reasoning for control through the lens of variational inference.

View source

Similar papers

Book Open access Aug 2026

InfRL: Inference-time Reinforcement Learning for Research Idea Optimization

InfRL (Inference-time Reinforcement Learning) offers a practical and domain-agnostic approach to harness reinforcement learning during inference, bridging the gap between static prompting and computationally intensive parameter-level fine-tuning.

Sikun Guo, Amir Hassan Shariatmadari, Jiuqi Wang et al. · 0 citations
Jul 2026

In-Context Learning as Implicit Policy Gradient

It is shown that score-conditioned In-Context Learning (ICL) admits a structural correspondence to policy gradient optimization, and an exact upper bound on the distribution shift induced by a bounded attention update is derived, yielding a trust-region-like analogy to KL-constrained policy optimization.

Masahiro Kaneko, Timothy Baldwin · 0 citations
Preprint Sep 2026

Large Discrete Policy: Advancing Explicit Behavior Modeling with Stochastic Iterative Scoring

Behavior policies are often formulated as continuous generative models, whose iterative denoising processes are expressive but difficult to interpret and prone to producing implausible actions. We propose the Large Discrete Policy (LDiP), a fully discrete behavior modeling framework that selects actions from a large vo...

Zhen-Xin Li, Nadine Chang, Xing-Long Sun et al. · 1 citation
Conference Sep 2026

Raise One and Infer Three: Toward Reasoning- and Memory-Augmented Diffusion Policy Generalization

This work proposes "raise one and infer three" diffusion policy (ROITDP), a novel approach that introduces two complementary mechanisms, including a reasoning mechanism built upon the Chain-of-Skill Noise Watermark, which enables temporally coherent multi-step reasoning throughout the diffusion process under distributi...

Yi-Hang Zhu, Yuxuan Wang, Tong Li et al. · 0 citations
Preprint Aug 2026

Steering Recurrent Reasoners at Inference Time with Readout Feedback

Readout Feedback (RoFB), a test-time intervention that converts intermediate predictions into token-wise pairwise coupling forces injected into the latent dynamics, is introduced, suggesting that closed-loop steering of latent dynamics can serve as a complementary inference-time control mechanism for recurrent reasonin...

Shunsuke Kamiya, Masanori Koyama, Seongcheol Jeong et al. · 0 citations
Jul 2026

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning

The World Critic Model is proposed, built on a lightweight LeJEPA architecture; WCM jointly predicts future latent state and estimates values, such that the critic's representation is explicitly trained to capture temporal dynamics rather than merely regress scalar returns.

Senyu Fei, Xiaopeng Yu, Siyin Wang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.