Skip to content

Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks

Jul 2026 · arXiv.org · Vol abs/2607.13305 · 0 citations · 57 references
Computer Science

TL;DR

The Visual Dependency Gap (VDG) is introduced, the difference in per-question correctness between original-video and black-screen conditions, which motivate VDG as a standard audit for whether video benchmarks measure visually grounded capability.

Abstract

Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption across twenty models spanning 2-78B parameters and ten architecture families. We introduce the Visual Dependency Gap (VDG), the difference in per-question correctness between original-video and black-screen conditions. Paired McNemar tests on MVBench show that accuracy and visual dependency are separable: models differ on original video (p = 0.0003) but not on black screens (p = 0.53). Across models, task-type rankings are stable: Attribute Perception is strongly visual, whereas Temporal Reasoning approaches the language-only baseline. A diagnostic ladder from black screen to single frame, shuffled frames, and original video reveals that frame diversity supplies most of the visual benefit, while temporal order contributes near-zero accuracy across sixteen open-weight models. An ablation from 0.5 to 24 FPS rules out sparse sampling as the cause. H.264 experiments further show that stable aggregate accuracy conceals bidirectional question-level answer flips. The diagnostic also generalizes to four API-accessed models, whose VDG values range from 0.025 to 0.315. These results motivate VDG as a standard audit for whether video benchmarks measure visually grounded capability. Code is available at https://github.com/JaeLee18/accuracy-without-grounding.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, ne...

Long-Yin Zhang, P. Mahendra, Chengwei Wei et al. · 1 citation
#computer vision Preprint Aug 2026

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

It is indicated that current VLMs can identify anomalies within individual frames but struggle to integrate information across frames to reason about temporal consistency, and TimeCatch provides a controlled benchmark for evaluating temporal grounding in vision-language models.

Marek Hradil, Danae Sánchez Villegas · 0 citations
#artificial intelligence Preprint Sep 2026

From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

Video generation has advanced to produce visually compelling and temporally coherent results. Yet, whether these models can genuinely think with video--executing symbolic rules, respecting physical laws, and pursuing intentional goals--remains an open question. Existing benchmarks only partially address this, often con...

Meng Luo, Yi-Chen Liu, Jia-Hao Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TempCloze: Can Video-LLMs Identify the Missing Middle?

Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors, so TempCloze is introduced, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs.

Wenqi Pei, Henry Hengyuan Zhao, Yi-Lai Liu et al. · 0 citations
#natural language process... Preprint Aug 2026

DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos

DVBench is introduced, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives, and decompose data video understanding into five dimensions, identifying two notable phenomena.

Bomiao Wang, Zekai Shao, Jiexiang Lan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.