Skip to content

ChronoStitch: Training-Free Composition of Visual KV Memories for Long-Horizon Temporal Reasoning

Jul 2026 · arXiv.org · Vol abs/2607.19547 · 0 citations · 12 references
Computer Science Engineering

TL;DR

ChronoStitch selectively recomputes a small fraction of high-deviation later-chunk visual tokens while allowing them to attend over the composed cache, improving event-ordering accuracy while running 3.3x faster than full joint re-prefilling.

Abstract

Long-video question answering requires a model to preserve visual evidence over time without repeatedly reprocessing the same video. A practical approach is to store the vision-language model's internal key-value (KV) cache for each video chunk and retrieve that state at query time. However, independently cached video chunks do not compose correctly: every chunk is prefilled from local rotary position zero, so naive concatenation collides temporal phases and removes the global order required for questions about what happened first, how often events occurred, or what changed across the video. This paper presents ChronoStitch, a training-free method for composing independently stored visual KV memories. The method first re-bases stored post-rotary keys onto a global three-axis multimodal RoPE coordinate system that preserves time, height, and width structure. We show why a one-dimensional scalar re-indexing is geometrically inconsistent for visual tokens because it turns spatial order within a frame into false temporal displacement. We then address the residual content gap left by positional repair: later chunks were originally encoded without attending to earlier chunks. ChronoStitch therefore selectively recomputes a small fraction of high-deviation later-chunk visual tokens while allowing them to attend over the composed cache. On Qwen2.5-VL-3B and the temporal split of TempCompass, ChronoStitch outperforms naive composition and position-only variants, improving event-ordering accuracy while running 3.3x faster than full joint re-prefilling.

View source

Similar papers

#computer vision Preprint Aug 2026

Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding

This work forms an episodic multi-key representation that enables precise retrieval of fine-grained memories through a simple key-matching mechanism, and introduces a neighbor filtering mechanism to capture broader semantic context without the massive computational overhead of global memory construction.

Yeeun Choi, Youngbeom Yoo, Joon-Young Lee et al. · 0 citations
Preprint Aug 2026

Addressable Memory for Video World Models

This work proposes WorldTrace, a training-free memory framework for long-horizon visual persistence that keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position.

Xindi Wu, Sven Elflein, James Lucas et al. · 1 citation
Preprint Aug 2026

Where Should Experience Live? Hierarchical Hebbian Memory for Continual Vision Transformers

Vision Transformers provide strong visual representations but typically rely on slowly updated parameters, limiting their ability to organize newly acquired information across different memory timescales. This work proposes \textit{Hierarchical Hebbian Memory}, a three-level memory architecture composed of rapid Workin...

Mohammed Yusuf Mujawar, Noorbakhsh Amiri Golilarz · 0 citations
Preprint Aug 2026

REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering

ReVEAL consistently outperforms both closed-source and open-source state-of-the-art methods across extensive experiments and shows that explicitly verifying evidence sufficiency, rather than stopping at semantic relevance, retrieves the decisive clues that prior methods miss and yields more reliable long-video reasonin...

C.J. Yan, Yang Zhou, Meixing Shi et al. · 1 citation
Preprint Aug 2026

Temporal Tree of Thought: Reasoning-Guided Visual Cue Search for Long-Video Understanding

The proposed Temporal Tree of Thought T^3, a training-free framework for adaptive coarse-to-fine long-video understanding, constructs a question-agnostic hierarchical temporal tree via recursive temporally constrained clustering, where each node represents a contiguous segment with an informative key frame.

Ziling Huang, Shin'ichi Satoh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.