Skip to content
Open access

Beyond VQA Accuracy: A Cross-Regime Diagnostic Evaluation of Symbol Consistency in Multimodal Large Language Models

2026 · IEEE Access · Vol 14, pp. 109309-109331 · 0 citations · 44 references
Computer Science

Abstract

Multimodal large language models (MLLMs) have achieved strong performance in general visual question answering, yet their visual faithfulness in handling discrete symbolic information remains insufficiently understood. Symbols such as numbers, time expressions, identifiers, license plates, and alphanumeric strings impose low semantic redundancy and strict character-level constraints, making overall VQA accuracy or isolated OCR-style evaluation inadequate for diagnosing model reliability. To address this gap, this paper proposes a cross-regime diagnostic framework for evaluating symbol consistency in MLLMs. Under a unified protocol, we evaluate five representative models, including BLIP-2, InstructBLIP, LLaVA, InternVL, and Qwen, across VQA, TextVQA, a self-constructed Symbol Subset, Regime 2a with uncontrolled generative symbol rendering, and Regime 2b with controlled clear-symbol grounding. We further introduce a 2a $\rightarrow 2$ b paired recovery analysis to distinguish rendering-sensitive errors caused by upstream symbol degradation from persistent errors that remain under clear visual evidence. Regime 2a is interpreted as an uncontrolled stress probe rather than a clean OCR benchmark, and its accuracy reflects both upstream rendering quality and downstream model reading. Experimental results show that general VQA accuracy can substantially overestimate symbol-centered reliability, especially for weaker models. Symbol consistency failures are not merely OCR recognition errors, but arise from the interaction of target-region binding, character-faithful transcription, answer completeness, and language-prior normalization. Although clear-symbol conditions improve stronger models, persistent failures remain in character-level grounding, target binding, and task following. This study separates symbol consistency from general VQA evaluation and reveals its multi-stage failure mechanisms, offering a reusable diagnostic perspective for reliability assessment and symbol-capability improvement in MLLMs.

Read PDF