Skip to content
Conference Open access

MagicBench: Diagnosing Visual Agency Loss and Semantic Dependency in Multimodal LLMs

2026 · Annual Meeting of the Association for Computational Linguistics · pp. 28493-28511 · 0 citations · 22 references
Computer Science

TL;DR

These findings suggest MLLMs function as language-guided passive observers advocating for perceptually-independent architectures that decouple sensory agency from linguistic dominance, and Causal interventions via spatial prompting and signal magnification provide evidence that internal reasoning remains functional, supporting the interpretation of a perceptual access bottleneck.

Abstract

Multimodal Large Language Models typically assume linguistic context invariably enhances visual understanding. We study this assumption in semantic adversarial scenarios, specifically magic tricks, where narration deliberately diverges from physical reality. We introduce MagicBench, a diagnostic benchmark of 402 videos for evaluating MLLMs under hierarchical linguistic interference, together with a Physical Constraint Set (PCS) protocol for assessing adherence to physical laws. Evaluation uncovers a Semantic Dependency Paradox: (1) Semantic anchoring : Entity nouns act as anchors aiding localization, paradoxically boosting performance despite false predicates. (2) Visual Agency Loss : In semantic vacuums, multi-modal performance collapses 12.4% ( p < 0 . 01 ) below the vision-only capability probe . This gap persists under symmetric prompting, suggesting a form of functional perception suppression in which autonomous visual search may be under-utilized in multimodal settings without linguistic triggers. Causal interventions via spatial prompting and signal magnification provide evidence that internal reasoning remains functional, supporting the interpretation of a perceptual access bottleneck. Our findings suggest MLLMs function as language-guided passive observers , advocating for perceptually-independent architectures that decouple sensory agency from linguistic dominance. Code and dataset are available at https://github.

Read PDF

Similar papers

Preprint Aug 2026

Why Vision Fails as a Universal Bridge: Rectifying Modality Asynchrony in Multilingual MLLMs

Through rigorous mechanistic analysis, this work identifies the Ghost Anchor phenomenon: a temporal modality asynchrony where linguistic translation to the English semantic manifold completes in early layers, while visual semanticization remains immature.

Yihang Du, Juhao Liang, Zheng-Zhao Lai et al. · 0 citations
#artificial intelligence Preprint Aug 2026

EviAnchor: Mitigating Hallucinations in Large Vision-Language Models via Regional Visual Evidence Compensation

EviAnchor is proposed, a training-free and single-branch inference framework that preserves and reactivates visual evidence throughout generation in large vision-language models and demonstrates consistent improvements in visual grounding.

Sihang Jia, Shuliang Liu, Song-Bo Yang et al. · 0 citations
Preprint Aug 2026

Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models

This paper develops a comprehensive where-what-how taxonomy that characterizes where situational illusions occur, what targets they take, and how they arise, and introduces MSIBench, a benchmark designed to assess the discrimination, understanding, and reasoning capabilities of MLLMs under situational illusions.

Zhi-Ming Yang, Zhuoxi Xiong, Dong-Lin Zhou et al. · 0 citations
#computer vision Preprint Sep 2026

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-based decisions. With this purpose in mind, we investigate cross-modal information flow in a video-base...

Davide Testa, Hugh Mee Wong, Alessandro Lenci et al. · 0 citations
Jul 2026

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

For the coarse attributes the authors study, MLLMs encode the visual evidence but cannot reliably control their reliance on it, indicating that for the coarse attributes they study, MLLMs cannot reliably control their reliance on it.

Jiaang Li, Chengzu Li, Zhaochong An et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.