Skip to content
Preprint

Improving Spatial-Temporal Reasoning in Video-Language Models with Structured Video Prompting

Aug 2026 · 0 citations · 21 references
Computer Science

TL;DR

Structured video prompting is introduced, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding.

Abstract

Video-language models (VLMs) remain brittle on tasks that require tracking events over time and grounding answers in specific spatial regions. We propose that part of this limitation can be addressed through better organization of visual evidence at inference time. We introduce structured video prompting, a training-free inference-time method that augments the input video with lightweight spatial structure and temporal structure, providing explicit anchors for organizing evidence across space and time without changing model weights or decoding and without altering the question prompt in the main comparison. We evaluate this approach on two complementary video benchmarks and two open video-language models. Across these settings, structured inputs improve performance in several cases, with gains varying by model and task. Our findings suggest that some failures of VLMs arise not only from reasoning capacity, but also from how video evidence is presented at inference time. These results highlight structured video prompting as a simple and practical direction for improving video understanding.

View source

Similar papers

Preprint Aug 2026

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While Thinking-with-Videos enables active temporal perception and Deep Research supports multi-step information seeking, the two capabilities are...

Wenqi Liu, Shi-Jie Ma, Yun-Xiao Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

Kairos is introduced, a video dataset for video-language modeling with time-resolved annotations that supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation.

Ruibo Ming, Lei Sun, De-Heng Zhang et al. · 0 citations
Preprint Sep 2026

The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs

VT-Contrast supervises selected late-layer last-frame video tokens, where temporal information is expected to be integrated before language generation, and contrasts order-preserving views with same-video reordered counterfactuals graded by Kendall tau distance.

Yu-Meng Shi, Quan-Yu Long, Yin Wu et al. · 0 citations
Preprint Aug 2026

Temporal Tree of Thought: Reasoning-Guided Visual Cue Search for Long-Video Understanding

The proposed Temporal Tree of Thought T^3, a training-free framework for adaptive coarse-to-fine long-video understanding, constructs a question-agnostic hierarchical temporal tree via recursive temporally constrained clustering, where each node represents a contiguous segment with an informative key frame.

Ziling Huang, Shin'ichi Satoh · 0 citations
#computer vision Review Sep 2026

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count, and organize methods by the pipeline stage at which they act.

Killian Steunou, Yannis Tevissen, M. El Yacoubi · 0 citations
#artificial intelligence Preprint Sep 2026

From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

This work proposes Vid-PRE (Video Prompt Reasoner and Enhancer), a model-agnostic prompt rewriter that offloads the cognitive burden of reasoning to a dedicated VLM, and introduces VWG-Bench, a comprehensive benchmark spanning 9 reasoning dimensions and 38 fine-grained tasks.

Meng Luo, Yi-Chen Liu, Jia-Hao Wang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.