2026· Annual Meeting of the Association for Computational Linguistics· pp. 28493-28511· 0 citations· 22 references
Computer Science
TL;DR
These findings suggest MLLMs function as language-guided passive observers advocating for perceptually-independent architectures that decouple sensory agency from linguistic dominance, and Causal interventions via spatial prompting and signal magnification provide evidence that internal reasoning remains functional, supporting the interpretation of a perceptual access bottleneck.
Abstract
Multimodal Large Language Models typically assume linguistic context invariably enhances visual understanding. We study this assumption in semantic adversarial scenarios, specifically magic tricks, where narration deliberately diverges from physical reality. We introduce MagicBench, a diagnostic benchmark of 402 videos for evaluating MLLMs under hierarchical linguistic interference, together with a Physical Constraint Set (PCS) protocol for assessing adherence to physical laws. Evaluation uncovers a Semantic Dependency Paradox: (1) Semantic anchoring : Entity nouns act as anchors aiding localization, paradoxically boosting performance despite false predicates. (2) Visual Agency Loss : In semantic vacuums, multi-modal performance collapses 12.4% ( p < 0 . 01 ) below the vision-only capability probe . This gap persists under symmetric prompting, suggesting a form of functional perception suppression in which autonomous visual search may be under-utilized in multimodal settings without linguistic triggers. Causal interventions via spatial prompting and signal magnification provide evidence that internal reasoning remains functional, supporting the interpretation of a perceptual access bottleneck. Our findings suggest MLLMs function as language-guided passive observers , advocating for perceptually-independent architectures that decouple sensory agency from linguistic dominance. Code and dataset are available at https://github.
Through rigorous mechanistic analysis, this work identifies the Ghost Anchor phenomenon: a temporal modality asynchrony where linguistic translation to the English semantic manifold completes in early layers, while visual semanticization remains immature.
Yihang Du, Juhao Liang, Zheng-Zhao Lai et al.· 0 citations
EviAnchor is proposed, a training-free and single-branch inference framework that preserves and reactivates visual evidence throughout generation in large vision-language models and demonstrates consistent improvements in visual grounding.
Sihang Jia, Shuliang Liu, Song-Bo Yang et al.· 0 citations
This paper develops a comprehensive where-what-how taxonomy that characterizes where situational illusions occur, what targets they take, and how they arise, and introduces MSIBench, a benchmark designed to assess the discrimination, understanding, and reasoning capabilities of MLLMs under situational illusions.
Zhi-Ming Yang, Zhuoxi Xiong, Dong-Lin Zhou et al.· 0 citations
Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-based decisions. With this purpose in mind, we investigate cross-modal information flow in a video-base...
Davide Testa, Hugh Mee Wong, Alessandro Lenci et al.· 0 citations
For the coarse attributes the authors study, MLLMs encode the visual evidence but cannot reliably control their reliance on it, indicating that for the coarse attributes they study, MLLMs cannot reliably control their reliance on it.
Jiaang Li, Chengzu Li, Zhaochong An et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.