This work proposes WorldTrace, a training-free memory framework for long-horizon visual persistence that keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position.
Abstract
We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. However, we find that models can no longer reliably address stored content once rollouts extend beyond the training horizon, because temporal Rotary Positional Embeddings (RoPE) offsets then fall outside the range seen during training and the model struggles to retrieve the relevant visual information through attention. Moreover, naively compressing the cache in the RoPE-rotated space corrupts memory by averaging together incompatible positional phases. To address this, we propose WorldTrace, a training-free memory framework for long-horizon visual persistence. WorldTrace keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position. Within this addressable cache, we study two memory compression approaches: WorldTrace-Field compresses history for temporal coherence, while WorldTrace-Landmark stores verbatim scene traces at detected transitions for episodic recall. We further introduce LoopBench, a benchmark evaluating whether a compressed cache can reconstruct a previously visited scene after a long detour. WorldTrace-Field improves temporal consistency by +15.5%, and WorldTrace-Landmark improves episodic recall by +19.5% on LoopBench, extending visually persistent generation without retraining.
This work proposes to decouple the frequencies of memory updates and Memory Caching, and introduces two critical mechanisms: an auxiliary Memory Loss that forces persistent internalization of the scene, and a Memory Caching strategy that regularizes active weights against catastrophic drift.
Baback Elmieh, Lynn Tsai, Zeman Li et al.· arXiv.org· 0 citations
Dynamic Hub-and-Spoke Memory is proposed, a training-free framework that represents distant history as structured textual memory while preserving the recent frames as visual tokens for fine-grained perception in streaming video understanding.
Xinru Jiang, Lin Zhao, Xi Xiao et al.· 4 citations
StreamFlow is introduced, an efficient visual memory framework that enables dynamic, on-demand access to historical visual information and improves the visual attention score while reducing end-to-end latency and peak memory, enabling more visually grounded and efficient reasoning.
Muxin Fu, Yifan Zhang, Wentao Zhang et al.· 0 citations
ChronoStitch selectively recomputes a small fraction of high-deviation later-chunk visual tokens while allowing them to attend over the composed cache, improving event-ordering accuracy while running 3.3x faster than full joint re-prefilling.
Santiram Tiwari, Nishant Sinha, K. Kislay· arXiv.org· 0 citations
This work introduces a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence, regardless of actual context length.
Jiacong Xu, Hanwen Jiang, Zhixin Shu et al.· arXiv.org· 5 citations· ⚡1
Recently, many streaming video understanding methods have been proposed by constructing an external memory to store historical data for computational reduction. Most methods focus on optimizing the injection procedure of current data (write) and retrieving informative historical data (read) from memory, while overlooki...