ChronoStitch selectively recomputes a small fraction of high-deviation later-chunk visual tokens while allowing them to attend over the composed cache, improving event-ordering accuracy while running 3.3x faster than full joint re-prefilling.
Abstract
Long-video question answering requires a model to preserve visual evidence over time without repeatedly reprocessing the same video. A practical approach is to store the vision-language model's internal key-value (KV) cache for each video chunk and retrieve that state at query time. However, independently cached video chunks do not compose correctly: every chunk is prefilled from local rotary position zero, so naive concatenation collides temporal phases and removes the global order required for questions about what happened first, how often events occurred, or what changed across the video. This paper presents ChronoStitch, a training-free method for composing independently stored visual KV memories. The method first re-bases stored post-rotary keys onto a global three-axis multimodal RoPE coordinate system that preserves time, height, and width structure. We show why a one-dimensional scalar re-indexing is geometrically inconsistent for visual tokens because it turns spatial order within a frame into false temporal displacement. We then address the residual content gap left by positional repair: later chunks were originally encoded without attending to earlier chunks. ChronoStitch therefore selectively recomputes a small fraction of high-deviation later-chunk visual tokens while allowing them to attend over the composed cache. On Qwen2.5-VL-3B and the temporal split of TempCompass, ChronoStitch outperforms naive composition and position-only variants, improving event-ordering accuracy while running 3.3x faster than full joint re-prefilling.
This work forms an episodic multi-key representation that enables precise retrieval of fine-grained memories through a simple key-matching mechanism, and introduces a neighbor filtering mechanism to capture broader semantic context without the massive computational overhead of global memory construction.
Yeeun Choi, Youngbeom Yoo, Joon-Young Lee et al.· 0 citations
This work proposes WorldTrace, a training-free memory framework for long-horizon visual persistence that keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position.
Xindi Wu, Sven Elflein, James Lucas et al.· 1 citation
Vision Transformers provide strong visual representations but typically rely on slowly updated parameters, limiting their ability to organize newly acquired information across different memory timescales. This work proposes \textit{Hierarchical Hebbian Memory}, a three-level memory architecture composed of rapid Workin...
Mohammed Yusuf Mujawar, Noorbakhsh Amiri Golilarz· 0 citations
GROVE is introduced, a training-free framework that supports both behaviors with one memory grown causally from a continuous video stream, and achieves the best results among the compared methods.
Sitong Gong, Caixin Kang, Tianyu Yan et al.· 0 citations
ReVEAL consistently outperforms both closed-source and open-source state-of-the-art methods across extensive experiments and shows that explicitly verifying evidence sufficiency, rather than stopping at semantic relevance, retrieves the decisive clues that prior methods miss and yields more reliable long-video reasonin...
C.J. Yan, Yang Zhou, Meixing Shi et al.· 1 citation
The proposed Temporal Tree of Thought T^3, a training-free framework for adaptive coarse-to-fine long-video understanding, constructs a question-agnostic hierarchical temporal tree via recursive temporally constrained clustering, where each node represents a contiguous segment with an informative key frame.
Ziling Huang, Shin'ichi Satoh· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.