The Visual Dependency Gap (VDG) is introduced, the difference in per-question correctness between original-video and black-screen conditions, which motivate VDG as a standard audit for whether video benchmarks measure visually grounded capability.
Abstract
Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption across twenty models spanning 2-78B parameters and ten architecture families. We introduce the Visual Dependency Gap (VDG), the difference in per-question correctness between original-video and black-screen conditions. Paired McNemar tests on MVBench show that accuracy and visual dependency are separable: models differ on original video (p = 0.0003) but not on black screens (p = 0.53). Across models, task-type rankings are stable: Attribute Perception is strongly visual, whereas Temporal Reasoning approaches the language-only baseline. A diagnostic ladder from black screen to single frame, shuffled frames, and original video reveals that frame diversity supplies most of the visual benefit, while temporal order contributes near-zero accuracy across sixteen open-weight models. An ablation from 0.5 to 24 FPS rules out sparse sampling as the cause. H.264 experiments further show that stable aggregate accuracy conceals bidirectional question-level answer flips. The diagnostic also generalizes to four API-accessed models, whose VDG values range from 0.025 to 0.315. These results motivate VDG as a standard audit for whether video benchmarks measure visually grounded capability. Code is available at https://github.com/JaeLee18/accuracy-without-grounding.
VSI is not a universal best abstention signal; it is a sample-intrinsic indicator of vision-ignoring failure, best used as a conditional ensemble component.
Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, ne...
Long-Yin Zhang, P. Mahendra, Chengwei Wei et al.· 1 citation
It is indicated that current VLMs can identify anomalies within individual frames but struggle to integrate information across frames to reason about temporal consistency, and TimeCatch provides a controlled benchmark for evaluating temporal grounding in vision-language models.
Video generation has advanced to produce visually compelling and temporally coherent results. Yet, whether these models can genuinely think with video--executing symbolic rules, respecting physical laws, and pursuing intentional goals--remains an open question. Existing benchmarks only partially address this, often con...
Meng Luo, Yi-Chen Liu, Jia-Hao Wang et al.· 0 citations
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors, so TempCloze is introduced, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs.
Wenqi Pei, Henry Hengyuan Zhao, Yi-Lai Liu et al.· 0 citations
DVBench is introduced, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives, and decompose data video understanding into five dimensions, identifying two notable phenomena.
Bomiao Wang, Zekai Shao, Jiexiang Lan et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.