Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored. Annotating charts is a common yet challenging communicative task, requiring models to infer intended messages, interpret chart semantics, and place appropriate textual or graphical elements. To address this gap, we introduce ChartAnno, a benchmark for evaluating MLLMs on chart annotation generation. It contains 1,200 real-world charts with paired code and annotation instructions across three levels of instruction specificity. We evaluate 10 representative MLLMs under two primary input settings: (1) chart code alone and (2) both chart code and chart image, and further include a chart image-only ablation study. Results show that proprietary models remain stronger overall, although large-scale open-source models narrow the gap. More specific instructions improve annotation quality, while inferring abstract intent remains most difficult for current MLLMs. Providing chart images brings limited overall gains, with improvements mainly appearing in design-related metrics. These findings highlight chart annotation generation as a challenging task requiring semantic grounding and effective annotation design. Code and data will be released in a future version.
This work evaluates 10 state-of-the-art MLLMs and examines three factors that influence performance: reasoning patterns, auxiliary tools, and robustness to image perturbations, showing that MLLM accuracy decreases and varies substantially as computational complexity increases.
Ziyan Xiao, Yinghao Zhu, Wen-Ting Zhang et al.· 1 citation
Charts play a key role in scientific research, offering a concise and visual way to present complex data. For Multimodal Large Language Models (MLLMs), the ability to comprehend charts is critical, as it requires both visual perception and reasoning that bridges graphical and textual information. However, existing char...
Tan Yue, Rui Mao, Xuzhao Shi et al.· Proceedings of the 32nd ACM...· 1 citation
Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to...
Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al.· 1 citation
DVBench is introduced, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives, and decompose data video understanding into five dimensions, identifying two notable phenomena.
Bomiao Wang, Zekai Shao, Jiexiang Lan et al.· 0 citations
As multimodal large language models (MLLMs) support a growing range of input modalities, increasing work explores how to incorporate rough sketches to convey user intent. For annotated chart generation, it remains unclear what annotation sketches people provide and when such visual input helps MLLMs generate more usefu...
Yoon-Gu Oh, Seon Gyeom Kim, J. Choi et al.· 0 citations
This work introduces Edit2TikZ, a comprehensive benchmark for scientific figure editing tasks, featuring 1,548 diverse and high-quality samples, and constructs a human-aligned evaluation framework to measure whether a requested edit is completed while irrelevant content is preserved.
Zong-Yun Zhang, Jiacheng Ruan, Xian Gao et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 2, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents create expressive charts from compact, human-editable specifications. The post Flint: A visualization language for the AI era appeared first on Microsoft Research.
Able to defeat top-ranked human players and more efficient than other models, the new system could help decision-makers in military maneuvers or business negotiations.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.