Skip to content

FOLIO: Focused Semantic Memory for Streaming Video Understanding

Jul 2026 · arXiv.org · Vol abs/2607.13298 · 1 citation · ⚡ 1 influential · 80 references
Computer Science

TL;DR

FOLIO is introduced, a training-free focused semantic memory system that records important parts of the stream in higher detail while keeping surrounding context compactly and substantially reducing the cost of maintaining streaming memory by reserving detailed records for focused entities and storing surrounding context compactly.

Abstract

In online streaming video understanding, a video stream continues to arrive and queries may be issued at any time. Because streaming frames grow without bound, the system must continuously compress and retain information from the observed video prefix while future frames and future queries remain unknown. The core challenge is deciding what information to retain and how to organize the maintained history: as this history grows with the stream, memory cost increases and many redundant visual details are retained, whereas later queries often depend on specific entities, actions, and their temporal changes. To address this challenge, we introduce FOLIO, a training-free focused semantic memory system that records important parts of the stream in higher detail while keeping surrounding context compact. As the stream arrives, FOLIO updates memory at the segment level, guided by a dynamic focus state, combining a short-term visual buffer with a long-term semantic memory organized around observed entities and linked to a visual-evidence cache. At query time, lightweight hybrid retrieval combines direct matching over the structured memory with semantic query expansion. FOLIO achieves state-of-the-art performance, reaching 82.0/69.1 Perception/Backward accuracy on OVO-Bench with Qwen3-VL-8B and 74.5 overall accuracy on StreamingBench, while substantially reducing the cost of maintaining streaming memory by reserving detailed records for focused entities and storing surrounding context compactly.

View source

Similar papers

Preprint Aug 2026

StreamScout: Learning When to Look Deeper for Streaming Video Understanding

Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily focus on what to retain in a bounded memory, yet access that memory using the same fixed-cost procedure for every query, despite substantial variation in the evidence re...

Ce Zhang, Jing Bi, Jinxi He et al. · 2 citations
Preprint Aug 2026

Dynamic Hub-and-Spoke Memory for Streaming Video Understanding

Dynamic Hub-and-Spoke Memory is proposed, a training-free framework that represents distant history as structured textual memory while preserving the recent frames as visual tokens for fine-grained perception in streaming video understanding.

Xinru Jiang, Lin Zhao, Xi Xiao et al. · 4 citations
Preprint Aug 2026

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

StreamFlow is introduced, an efficient visual memory framework that enables dynamic, on-demand access to historical visual information and improves the visual attention score while reducing end-to-end latency and peak memory, enabling more visually grounded and efficient reasoning.

Muxin Fu, Yifan Zhang, Wentao Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

PREM: Prefix-Steered Recurrent Memory for Long-Video Understanding

Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens, or alter internal key-value (KV) caches. We introduce Prefix-Steered Recurrent Memory (PREM), a memory-token-free framework for frozen vision-language models (VLMs). PR...

Si-Ru Zhong, Qiong-Yan Wang, Xiao-Hui Lv et al. · 0 citations
Preprint Aug 2026

RECAP-Forcing: Retaining Content Appearances for Long Video Generation

This work proposes RECAP-Forcing, a training-free inference method with no additional learnable parameters that consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.

Haiyang Xu, Zheng Ding, Zhuowen Tu · 0 citations
Preprint Aug 2026

StreamEMS: Streaming Video Understanding with Self-Evolving Memory Scheme for Vision-Language Models

Recently, many streaming video understanding methods have been proposed by constructing an external memory to store historical data for computational reduction. Most methods focus on optimizing the injection procedure of current data (write) and retrieving informative historical data (read) from memory, while overlooki...

Yu-Xing Liu, Peiqin Zhuang, Ya-Li Wang · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.