Skip to content

ChartAnno: Evaluating MLLMs for Chart Annotation Generation

Aug 2026 · 0 citations
Computer Science

Abstract

Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored. Annotating charts is a common yet challenging communicative task, requiring models to infer intended messages, interpret chart semantics, and place appropriate textual or graphical elements. To address this gap, we introduce ChartAnno, a benchmark for evaluating MLLMs on chart annotation generation. It contains 1,200 real-world charts with paired code and annotation instructions across three levels of instruction specificity. We evaluate 10 representative MLLMs under two primary input settings: (1) chart code alone and (2) both chart code and chart image, and further include a chart image-only ablation study. Results show that proprietary models remain stronger overall, although large-scale open-source models narrow the gap. More specific instructions improve annotation quality, while inferring abstract intent remains most difficult for current MLLMs. Providing chart images brings limited overall gains, with improvements mainly appearing in design-related metrics. These findings highlight chart annotation generation as a challenging task requiring semantic grounding and effective annotation design. Code and data will be released in a future version.

View source

Similar papers

Preprint Aug 2026

LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning

This work evaluates 10 state-of-the-art MLLMs and examines three factors that influence performance: reasoning patterns, auxiliary tools, and robustness to image perturbations, showing that MLLM accuracy decreases and varies substantially as computational complexity increases.

Ziyan Xiao, Yinghao Zhu, Wen-Ting Zhang et al. · 1 citation
Book Open access Aug 2026

SciChart: Visual Question Answering and Reasoning for Scientific Spectral Chart

Charts play a key role in scientific research, offering a concise and visual way to present complex data. For Multimodal Large Language Models (MLLMs), the ability to comprehend charts is critical, as it requires both visual perception and reasoning that bridges graphical and textual information. However, existing char...

Tan Yue, Rui Mao, Xuzhao Shi et al. · 1 citation
#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to...

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation
#natural language process... Preprint Aug 2026

DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos

DVBench is introduced, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives, and decompose data video understanding into five dimensions, identifying two notable phenomena.

Bomiao Wang, Zekai Shao, Jiexiang Lan et al. · 0 citations
#human-computer interacti... Preprint Sep 2026

AnnoSketch: Evaluating and Collecting Human Sketches for MLLM-assisted Chart Annotation

As multimodal large language models (MLLMs) support a growing range of input modalities, increasing work explores how to incorporate rough sketches to convey user intent. For annotated chart generation, it remains unclear what annotation sketches people provide and when such visual input helps MLLMs generate more usefu...

Yoon-Gu Oh, Seon Gyeom Kim, J. Choi et al. · 0 citations
Preprint Aug 2026

Edit2TikZ: A Comprehensive and Challenging Benchmark for Scientific Figure Editing with TikZ

This work introduces Edit2TikZ, a comprehensive benchmark for scientific figure editing tasks, featuring 1,548 diverse and high-quality samples, and constructs a human-aligned evaluation framework to measure whether a requested edit is completed while irrelevant content is preserved.

Zong-Yun Zhang, Jiacheng Ruan, Xian Gao et al. · 0 citations

Related blog posts

Microsoft Research Blog Jul 8, 2026

Flint: A visualization language for the AI era

Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flint: A visualization language for the AI era appeared first on Microsoft Research.

MIT News · Artificial Intelligence Sep 30, 2026

This game-playing AI is the new champ at Stratego

Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.