Skip to content

Attention-Guided Saliency Maps for Interpreting Visualization Literacy in VLMs

Jul 2026 · arXiv.org · Vol abs/2607.16105 · 0 citations · 26 references
Computer Science

TL;DR

A lightweight, diagnostic saliency map method tailored for text generation over images using transformer models, the current state-of-the-art models in visualization interpretation, is introduced.

Abstract

Understanding how vision-language models (VLMs) interpret data visualizations remains an open problem, and is increasingly important as these models are used for analytical tasks where reliable reasoning is essential. We introduce a lightweight, diagnostic saliency map method tailored for text generation over images using transformer models, the current state-of-the-art models in visualization interpretation. Our approach aggregates the language model's attention over the visual tokens across all heads and layers, then maps this attention back onto the vision encoder's patch grid to localise it over the image, producing a direct correspondence between each generated answer token and the image regions it attended to. This yields fast, gradient-free saliency maps that expose how VLMs allocate focus across visual elements during answer generation, enabling inspection of whether model attention aligns with semantically relevant components. We evaluate our approach using a deletion metric which validates the causal faithfulness of our saliency maps to the model's behavior.

View source

Similar papers

#computer vision Preprint Aug 2026

ReVA: A Region-Aware Visual Assistant for Visually Grounded Question Answering

ReVA is proposed, a region-aware VQA model that employs a frozen CLIP ViT-L/14 Vision Transformer (ViT) and a Qwen2.5-7B-Instruct large language model (LLM) connected through a dual bridge that aligns both whole-image and region-level representations with the LLM's embedding space.

A. Senthil · 0 citations
Preprint Aug 2026

Where To Look? : Causal Tracing of Vision Encoders in VLM

This work observes that highly causal vision tokens often lie outside the target region, and extends the analysis to larger vision-language models, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations.

S. NarenKumar, T. Bhatt, Mayank Singh · 0 citations
#small language model Preprint Aug 2026

Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

This work introduces Visual Retrieval Heads (VRHs), a small subset of attention heads that are causally responsible for grounding text descriptions to image regions, and shows that scoring attention from output prediction tokens with a sum over the ground-truth referent region most reliably identifies causal heads.

Chanho Park, Daehyeon Choi, Jihyun Lee et al. · 0 citations
Preprint Aug 2026

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model, is proposed, endowing models with native fine-grained region description and flexible reasoning capabilities.

Chang-Jiang Jiang, Qian-Nian Zhao, Lei Xin et al. · 2 citations
Jul 2026

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models and parses them into structured predictions compatible with original metrics, is proposed and revealed, revealing complementary strengths of pixel-space expression and t...

Xu Wang, Kaixiang Yao, Miao Pan et al. · 1 citation
Jul 2026

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models

For the coarse attributes the authors study, MLLMs encode the visual evidence but cannot reliably control their reliance on it, indicating that for the coarse attributes they study, MLLMs cannot reliably control their reliance on it.

Jiaang Li, Chengzu Li, Zhaochong An et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.