Skip to content

InSight: A Benchmark for Agentic Claim Verification in Interactive Visualizations

Sep 2026 · 0 citations · 34 references
Computer Science

TL;DR

This paper introduces InSight, a benchmark for agentic claim verification over interactive visualizations and evaluates state-of-the-art models, revealing that interactive verification remains a non-trivial challenge.

Abstract

Vision Language Models have demonstrated remarkable proficiency in interpreting static visual artifacts, but modern data analysis is inherently dynamic, requiring the active interrogation of interactive environments. Existing benchmarks are predominantly constrained to static imagery and one-shot question answering and fail to capture the epistemic demands of this domain, where evidence is frequently occluded, distributed across linked views, or conditionally revealed through user agency. In this paper, we introduce InSight, a benchmark for agentic claim verification over interactive visualizations. The dataset consists of 21,349 claims derived from human-authored analytical narratives and grounded in fully interactive web-based environments. Agents must navigate these environments to determine whether a natural language claim is supported, refuted or not verifiable given the available evidence. Unlike traditional evaluations, InSight treats interaction traces as intrinsic proxies for reasoning, enabling a rigorous audit of how models seek and synthesize visual evidence. We evaluate state-of-the-art models, revealing that interactive verification remains a non-trivial challenge. We release InSight at https://github.com/maevehutch/insight.

View source

Similar papers

Preprint Aug 2026

SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding

SAGE is an evidence-grounded multi-agent framework that reformulates Chinese ancient document understanding as evidence-grounded inference rather than direct answer generation, highlighting the importance of structured, evidence-grounded inference beyond model scaling.

Yu-Chuan Wu, Xuan Luo, Yinglian Zhu et al. · 0 citations
Open access Sep 2026

Bridging trust and performance in intelligent systems: Hybrid explainable AI approaches for interpreting large language models

A hybrid explainability framework that integrates saliency-based attribution, causal reasoning, and user-centered visualization into a unified, efficiency-aware pipeline is introduced, positioning hybrid XAI as a pathway toward responsible LLM adoption.

A. P, Tamije Selvy P · 0 citations
Preprint Aug 2026

VizAnchor: Decoding Manipulation Intent from Tampering Visualizations via Dual-Anchor Reasoning

Data visualizations are widely used for communicating information, but they are also vulnerable to intentional manipulations that induce misleading interpretations. Existing methods focus on locating tampered regions or recovering hidden information, without explaining how the visualization has been manipulated or why...

Xiaotian Zhang, Huayuan Ye, Haiyang Zhang et al. · 0 citations

Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization

Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential basis of these interpretations is often unclear. We present an exploratory study of 102 visualizations from four sources, three MLLMs, and four input conditions that var...

I. Eliza, Md Dilshadur Rahman · 0 citations
#natural language process... Preprint Aug 2026

Beyond Static Charts: Can Language and Vision Language Models Generate Interactive Data Visualization Interfaces?

A structured multi stage interface generation framework that decomposes the task into visualization design representation, generation of multiple interface candidates, constraint-aware critique, and self-refinement is proposed, demonstrating a practical path toward more reliable language-driven interactive visualizatio...

Mizanur Rahman, Aaryaman Kartha, Enamul Hoque Prince · 0 citations
Jul 2026

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.

Siyu Yan, Zhuoran Yan, Haiying Xu et al. · 0 citations

Related blog posts

Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.