Structured video prompting is introduced, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding.
Abstract
Video-language models (VLMs) remain brittle on tasks that require tracking events over time and grounding answers in specific spatial regions. We propose that part of this limitation can be addressed through better organization of visual evidence at inference time. We introduce structured video prompting, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding and without altering the question prompt in the main comparison. We evaluate this approach on two complementary video benchmarks and two open video-language models. Across these settings, structured inputs improve performance in several cases, with gains varying by model and task. Our findings suggest that some failures of VLMs arise not only from reasoning capacity, but also from how video evidence is presented at inference time. These results highlight structured video prompting as a simple and practical direction for improving video understanding.
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step information seeking, the two capabilities are...
Wenqi Liu, Shi-Jie Ma, Yun-Xiao Wang et al.· 0 citations
Kairos is introduced, a video dataset for video-language modeling with time-resolved annotations that supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation.
Ruibo Ming, Lei Sun, De-Heng Zhang et al.· 0 citations
VT-Contrast supervises selected late-layer last-frame video tokens, where temporal information is expected to be integrated before language generation, and contrasts order-preserving views with same-video reordered counterfactuals graded by Kendall tau distance.
Yu-Meng Shi, Quan-Yu Long, Yin Wu et al.· 0 citations
The proposed Temporal Tree of Thought T^3, a training-free framework for adaptive coarse-to-fine long-video understanding, constructs a question-agnostic hierarchical temporal tree via recursive temporally constrained clustering, where each node represents a contiguous segment with an informative key frame.
This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count, and organize methods by the pipeline stage at which they act.
Killian Steunou, Yannis Tevissen, M. El Yacoubi· 0 citations
This work proposes Vid-PRE (Video Prompt Reasoner and Enhancer), a model-agnostic prompt rewriter that offloads the cognitive burden of reasoning to a dedicated VLM, and introduces VWG-Bench, a comprehensive benchmark spanning 9 reasoning dimensions and 38 fine-grained tasks.
Meng Luo, Yi-Chen Liu, Jia-Hao Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.