Skip to content

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video

Jul 2026 · arXiv.org · Vol abs/2607.11078 · 0 citations · 18 references
Computer Science

TL;DR

A diagnostic toolkit for auditing what benchmark scores for open-source Video-LLM models actually measure, and finds the accuracy does not come from character tracking.

Abstract

Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, in order, how their outfit changes across a full TV episode? Benchmarks increasingly score this kind of task, and the strongest open-source 7--8B models now reach 37--38% on InfiniBench's global appearance task, which asks exactly that. But does that score come from tracking the named character, or from something easier? We test this with a nine-condition diagnostic protocol applied to three architecturally distinct open-source Video-LLMs, with Gemini~2.5~Flash as a frontier reference, and find the accuracy does not come from character tracking. When we change the character named in the question to a different cast member, leaving the video and answer options untouched, the models change their answer only 4--31% of the time, so they are largely ignoring who the question asks about. Breaking that test down by the gender of the swapped name shows why: the models react more when the name is changed to a different-gender character than to a same-gender one (a 13--28 point gap), picking up coarse gender cues but unable to tell same-gender individuals apart. This shallow processing surfaces again when we drop the multiple-choice options and ask the same questions open-endedly: open-source accuracy drops 18--25 points, with none of 151 answers fully correct, versus a 12-point drop for Gemini. Further checks rule out the obvious innocent explanations, adding subtitles, using the most informative frames, or doubling the number of frames all leave character tracking unimproved, so the bottleneck is not how much video the model sees but how it ties that video to the person the question names. We release a diagnostic toolkit for auditing what such benchmark scores actually measure.

View source

Similar papers

Preprint Sep 2026

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

MotionBlind, a contrastive benchmark of self-recorded video for physically grounded motion, the variables a world model must predict, is introduced, a pair of near-identical clips that differ only in motion.

Dhairya Bhatia, Bishoy M. Galoaa, Oliver Fritsche et al. · 0 citations
#computer vision Review Sep 2026

Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs

Video understanding has rapidly evolved toward video large language models (VideoLLMs): systems that couple video representations with pretrained large language models and condition generation on a textual prompt. Their strong performance on captioning, question answering, retrieval and temporal grounding comes at a co...

Killian Steunou, Yannis Tevissen, M. E. El Yacoubi · 0 citations

Position: Video LLMs Must Not Ignore the Pixel Dynamics in Plain Sight

It is argued that recent progress in video understanding is measured by benchmarks and protocols that can be solved without reliably perceiving spatiotemporal evidence, rewarding language-driven plausibility over video-grounded inference.

Shayda Moezzi, Umer Saleem, Andong Deng et al. · 1 citation
Preprint Sep 2026

The Shape of Time: Video-Token Contrast for Temporal Understanding in VideoLMs

Seeing frames in order does not mean representing time. Modern VideoLMs receive ordered video streams, yet their main supervision acts on generated text rather than video-token representations where event dynamics should first emerge. This mismatch allows models to learn temporal answers from shortcuts such as objects,...

Yu-Meng Shi, Quanyu Long, Yin Wu et al. · 0 citations
Preprint Aug 2026

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

This work conducts a large-scale evaluation of more than 20 recent MLLMs and shows that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based o...

Hongbo Liu, Peixian Chen, Siyuan Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.