Skip to content
Preprint

Addressable Memory for Video World Models

Aug 2026 · 1 citation · 120 references
Computer Science

TL;DR

This work proposes WorldTrace, a training-free memory framework for long-horizon visual persistence that keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position.

Abstract

We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. However, we find that models can no longer reliably address stored content once rollouts extend beyond the training horizon, because temporal Rotary Positional Embeddings (RoPE) offsets then fall outside the range seen during training and the model struggles to retrieve the relevant visual information through attention. Moreover, naively compressing the cache in the RoPE-rotated space corrupts memory by averaging together incompatible positional phases. To address this, we propose WorldTrace, a training-free memory framework for long-horizon visual persistence. WorldTrace keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position. Within this addressable cache, we study two memory compression approaches: WorldTrace-Field compresses history for temporal coherence, while WorldTrace-Landmark stores verbatim scene traces at detected transitions for episodic recall. We further introduce LoopBench, a benchmark evaluating whether a compressed cache can reconstruct a previously visited scene after a long detour. WorldTrace-Field improves temporal consistency by +15.5%, and WorldTrace-Landmark improves episodic recall by +19.5% on LoopBench, extending visually persistent generation without retraining.

View source

Similar papers

Jul 2026

Online Neural Space Time Memory for Dynamic Novel View Synthesis

This work proposes to decouple the frequencies of memory updates and Memory Caching, and introduces two critical mechanisms: an auxiliary Memory Loss that forces persistent internalization of the scene, and a Memory Caching strategy that regularizes active weights against catastrophic drift.

Baback Elmieh, Lynn Tsai, Zeman Li et al. · 0 citations
Preprint Aug 2026

Dynamic Hub-and-Spoke Memory for Streaming Video Understanding

Dynamic Hub-and-Spoke Memory is proposed, a training-free framework that represents distant history as structured textual memory while preserving the recent frames as visual tokens for fine-grained perception in streaming video understanding.

Xinru Jiang, Lin Zhao, Xi Xiao et al. · 4 citations
Preprint Aug 2026

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

StreamFlow is introduced, an efficient visual memory framework that enables dynamic, on-demand access to historical visual information and improves the visual attention score while reducing end-to-end latency and peak memory, enabling more visually grounded and efficient reasoning.

Muxin Fu, Yifan Zhang, Wentao Zhang et al. · 0 citations
Jul 2026

Wonder: Video World Model Done Better

This work introduces a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence, regardless of actual context length.

Jiacong Xu, Hanwen Jiang, Zhixin Shu et al. · 5 citations · ⚡1
Preprint Aug 2026

StreamEMS: Streaming Video Understanding with Self-Evolving Memory Scheme for Vision-Language Models

Recently, many streaming video understanding methods have been proposed by constructing an external memory to store historical data for computational reduction. Most methods focus on optimizing the injection procedure of current data (write) and retrieving informative historical data (read) from memory, while overlooki...

Yu-Xing Liu, Peiqin Zhuang, Ya-Li Wang · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.