Skip to content

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

Jun 2026 · arXiv.org · Vol abs/2606.17188 · 0 citations · 20 references
Computer Science

TL;DR

This work introduces PuMVR (Punjabi Multimodal Visual Reasoning), a benchmark of 1,000 strictly parallel image-text instances across Punjabi's three active scripts and proposes the Script Consistency Rate (SCR), which falls as low as 24.8% on this benchmark, as a mandatory metric for script-agnostic evaluation to ensure equitable AI access.

Abstract

Current multilingual evaluations for Vision-Language Models (VLMs) assume a one-to-one mapping between language and orthography, overlooking billions of users of multi-script languages. We introduce PuMVR (Punjabi Multimodal Visual Reasoning), a benchmark of 1,000 strictly parallel image-text instances across Punjabi's three active scripts: Gurmukhi, Shahmukhi, and Roman. Evaluating 10 state-of-the-art VLMs, we expose a substantial and systematic Script Gap. Models frequently solve visual tasks in one script while failing identical tasks in another, with accuracy deltas reaching 16%. Crucially, visual input boosts absolute performance uniformly yet does not close the orthographic gap. Furthermore, cross-script in-context transfer is highly brittle, exposing script-locked knowledge representation. Supported by McNemar tests across all script pairs, our findings demonstrate that current"multilingual"VLMs are not truly multi-script. We propose the Script Consistency Rate (SCR), which falls as low as 24.8% on our benchmark, as a mandatory metric for script-agnostic evaluation to ensure equitable AI access. Data and code are available at: https://github.com/prabhjotschugh/Not-Truly-Multilingual-PuMVR.

View source

Similar papers

Output Language Confusion under Multilingual Prompt Contamination

Standard factual benchmarks assume clean monolingual prompts and exact-match scoring, two assumptions that break simultaneously in real-world multilingual deployment, from retrieval-augmented generation pipelines returning mixed-language passages to users pasting multilingual web content. We introduce Multilingual Dist...

Riju Marwah, Ritvik Garimella, Khusham Bansal et al. · 0 citations
Preprint Sep 2026

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

This work constructs TextMuSS-10M, a large-scale synthetic scene text dataset spanning 10 scripts and 229 languages and proposes ScriptMoE, a script-aware Mixture-of-Experts (MoE) architecture that achieves the highest accuracy and is simpler than per-language experts, lighter than VLMs, and more accurate than both.

Xing-Song Ye, Yong-Kun Du, Jia-Xin Zhang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

VISTA-Bench: Benchmarking Multilingual Image Translation with Image-Specific Rubrics

Image translation is a fundamental capability of multimodal models for multilingual applications, requiring visual understanding and meaning preservation across languages. However, existing benchmarks have limited language coverage and often lack explicit image-specific evaluation criteria, making it difficult to compr...

Bo Lv, Mao Zheng, Zheng Li et al. · 0 citations
Preprint Aug 2026

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual performance gaps and language-specific effects insufficiently explored. To fill this gap, we introduce LingT2I, a benchmark covering 10 widely u...

Si-Cheng Zhang, Zhong-Hao Yan, Bin-Zhu Xie et al. · 0 citations
#artificial intelligence Preprint Oct 2026

Cross-Lingual Alignment for Decoder-Only Models using MoE Routers

Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, higher cross-lingual rep...

Lucas Bandarkar, C. Peng, A. H. Ahmed et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.