Skip to content
Preprint

VideoHarness-RSI: Recursive Harness Self-Improvement for Long-Video Understanding with Frozen Vision-Language Models

Aug 2026 · 0 citations · 22 references
Computer Science

TL;DR

VideoHarness-RSI is introduced, a controlled framework that recursively searches executable context constructors around a frozen VLM while keeping the answering model and interface fixed and establish executable context construction as a distinct optimization layer and provide an auditable baseline for studying harness discovery, transfer, and efficiency around frozen VLMs.

Abstract

Long-video understanding depends not only on the capability of a vision-language model (VLM), but also on how its limited context is constructed from a much longer video. Existing systems typically introduce hand-designed sampling, retrieval, memory, or agentic control strategies, making the context-construction program itself difficult to study as an independent optimization target. We introduce VideoHarness-RSI, a controlled framework that recursively searches executable context constructors around a frozen VLM while keeping the answering model and interface fixed. We study this baseline under complementary weak- and strong-initialization regimes. From a weak uniform constructor, recursive search progressively discovers more structured context-construction programs; from a stronger AKS harness, the same process further advances an already competitive hand-crafted frontier. The resulting harness retains its advantage under a matched cumulative visual-token control and transfers directly to additional long-video benchmarks without further search. Together, these results establish executable context construction as a distinct optimization layer and provide an auditable baseline for studying harness discovery, transfer, and efficiency around frozen VLMs.

View source

Similar papers

#machine learning Preprint Sep 2026

SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models

Long-horizon manipulation is partially observable: the information needed to choose the next action may appear only in observations from minutes earlier. Existing memory mechanisms: retrieval banks, learned compressors, recurrent states must decide what to keep from the past before knowing what a future decision will require. This was motivated by the assumption that minute-scale history is too large to process directly, which modern VLM backbones no longer make true. In this work, we introduce SimpleMemVLA, a VLA without a dedicated memory module. It keeps the sampled history intact and passes it to the backbone in the timestamped video format the backbone was pretrained to process; the hidden states of a generated sub-task then form the only channel from history to a standard flow-matching action head. Since consecutive decisions share most of their history, prefilling the shared prefix during action execution keeps latency close to a single-frame VLA. SimpleMemVLA sets a new state of the art on four memory benchmarks without cost on general-purpose control. Holding the backbone and training setup fixed, it outperforms retrieval, compression and recurrent-state mechanisms by a wide margin, and causal interventions confirm that the policy genuinely reads its history. Code available at https://github.com/wadeKeith/SimpleMemVLA

C. Yin, Wang Xu, Jun-Peng Yang et al. · 0 citations
Preprint Aug 2026

LongVU-TTT: Causal Test-Time Training for Visual Resampling in Long Video Understanding

LongVU-TTT is introduced, which inserts a convolutional Test-Time Training (TTT) resampler with causal fast-weight updates between the vision encoder and the LLM, and is stronger than attention- and fixed-state recurrent resamplers across three benchmarks.

Mahmoud Ahmed, Sameh Abdulah, Olatunji Ruwase et al. · 0 citations
Preprint Aug 2026

Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding

Clue-OPSD, a clue-privileged on-policy self-distillation framework for long-video understanding that uses clue intervals as privileged supervision without relying on ground-truth answer labels, while requiring no clue annotations or additional modules at inference time is introduced.

Kaishen Wang, Dong-Di Zhao, Yijun Liang et al. · 1 citation
#computer vision Review Sep 2026

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a computation and memory cost that grows with frame count and context length, limiting deployment in real-time, mobile and resource-constrained settings. This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count. We analyze bottlenecks across frame sampling, modality encoding, connector-level token reduction, and LLM prefilling and decoding. We organize methods by the pipeline stage at which they act, covering VideoLLMs developed since late 2022 together with earlier frame-sampling and vision-encoder mechanisms that remain components of current pipelines. We assemble literature-reported accuracy--cost comparisons under shared host models and input protocols wherever available, distinguish them from heterogeneous cross-paper evidence, and identify gaps in audiovisual efficiency and standardized evaluation. We maintain a repository at https://github.com/momentslab/awesome-efficient-videollm.

Killian Steunou, Yannis Tevissen, M. E. El Yacoubi · 0 citations
Jul 2026

Wonder: Video World Model Done Better

This work introduces a novel camera conditioning with a dense coordinate field whose renderings provide spatially aligned motion and orientation cues, allowing the model to interpret camera motion directly as visual evidence, regardless of actual context length.

Jiacong Xu, Hanwen Jiang, Zhixin Shu et al. · 4 citations · ⚡1
Preprint Aug 2026

Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs

VideoRouter (VR) is proposed that rethinks long-video understanding as coordinating complementary evidence views rather than selecting a single subset of frames, and introduces a verification-guided router to determine which view is better supported by the selected evidence and select the final answer.

Ziling Huang, Shin'ichi Satoh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.