Position: Video LLMs Must Not Ignore the Pixel Dynamics in Plain Sight
It is argued that recent progress in video understanding is measured by benchmarks and protocols that can be solved without reliably perceiving spatiotemporal evidence, rewarding language-driven plausibility over video-grounded inference.
Shayda Moezzi, Umer Saleem, Andong Deng et al.
· 0 citations