Skip to content

Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding

Sep 2026 · 0 citations · 41 references
Computer Science

TL;DR

This paper proposes a ``summarize before grounding''framework (named ``SumGround'') for long-video temporal grounding, and introduces query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries.

Abstract

Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs LVLMs to mine query-relevant evidence. Instead of dense frame sampling which incurs prohibitive training memory, previous reinforcement learning with verifiable rewards (RLVR) works typically utilize sparse sampling, which makes training feasible but may miss critical evidence. In this paper, we propose a ``summarize before grounding''framework (named ``SumGround'') for long-video temporal grounding. The key of SumGround is to perform query-guided chunk condensation to aggregate and retrieve query-relevant evidence. Specifically, we split the video into several chunks and perform two-level chunk condensation. First, we introduce query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries. Furthermore, we design an associative summary retrieval scheme to rank and select chunk summaries that are most likely to contain the event interval. Both query-guided latent summary and associative summary retrieval schemes are enabled by RLVR. To reduce memory consumption, we propose a length-aware gradient gating module to selectively stop gradient back-propagated to visual tokens. Extensive experiments demonstrate that SumGround performs favorably against previous state-of-the-art methods across multiple downstream datasets, with remarkable gains on long videos.

View source

Similar papers

#artificial intelligence Preprint Oct 2026

Enhancing Long-Video VLM Embeddings with Query-Aware Streaming Latent Reasoning

Long-video embedding requires capturing sparse query-relevant evidence under a limited visual-token budget. Uniform sampling can miss brief events in videos spanning minutes or hours, whereas encoding more frames in a single context increases memory and computation. We introduce \textbf{Query-Aware Streaming Latent Rea...

Hao-Zhe Chi, Song Jin, Yang Jin et al. · 0 citations
Preprint Aug 2026

Learning Compositional Spatio-Temporal Video Grounding with Synthetic Curriculum

This work builds a synthetic data engine that leverages a spatio-temporal scene graph as a difficulty measure and casts difficulty-controlled query synthesis as a constraint programming problem, producing difficulty-graded data for both evaluation and training and proposes CurrSTVG, a curriculum reinforcement learning...

Xing-Jian Wang, Shi-Jian Wang, Yi-Bo Wang et al. · 0 citations
Preprint Sep 2026

Beyond Similarity: Foundation Models as an Efficient Backbone for Training-Free Composed Video Retrieval

Composed video retrieval (CoVR) searches a gallery for the target video that realizes a natural-language modification of a source clip. However, at gallery scale, this creates a fundamental tension: compact embeddings enable efficient, reusable search but can miss the transient actions, state changes, and subtle constr...

Dmitry Demidov, Muhammad Zaigham Zaheer, Omkar Thawakar et al. · 0 citations
Review Sep 2026

Grounding with Confidence: Controllable Generative Video Temporal Grounding

Video temporal grounding supports applications such as video search, content review, and automated editing by localizing events described in natural language. Yet existing generative models typically output timestamps without explicit interval-level confidence scores to guide candidate selection. We separate candidate...

Jin-Hao Chen, Ben-Lei Cui, Rui-Jian Jia et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.