Skip to content
Book Open access

MVSCbench: A Benchmark for Symbolic and Cultural Understanding in Music Videos

Jul 2026 · Creativity & Cognition · 0 citations · 20 references
Computer Science

Abstract

This study introduces MVSCBench, the benchmark designed to evaluate multimodal AI models’ ability to interpret symbolic and cultural meanings in music videos. Using K-pop music video clips, we organize music video understanding into three stages: surface perception, symbolic interpretation, and cultural grounding. Based on this framework, we construct a dataset of 293 video clips and 3,516 annotations, covering space, artist, character symbolism, object symbolism, literary symbolism, social-historical symbolism, and fandom cultural symbols. The results show that current multimodal AI models perform well on surface perception tasks, but still face clear limitations in symbolic interpretation and cultural grounding. In particular, the models often struggle to correctly identify relevant visual cues and frequently produce hallucinated interpretations in categories involving deeper meanings. MVSCBench provides a new benchmark for evaluating symbolic and cultural understanding in music videos and offers a useful framework for future research.

Read PDF