Advancing Scientific Chart Understanding: The SCI-CQA Benchmark and Beyond.
Abstract
In real-world applications, current multimodal large models are often overestimated in their ability to understand scientific charts. To assess their true capabilities and identify key performance bottlenecks, we conducted an in-depth study on scientific chart understanding. Charts in scientific literature often feature complex visual elements, such as multi-plot figures, flowcharts, and structural diagrams. Evaluating multimodal models using such charts offers a more rigorous and comprehensive assessment of their visual and reasoning abilities. However, existing benchmarks fall short due to limited chart diversity, simplistic template-based questions, and inadequate evaluation protocols. To address these challenges, we first curated a high-quality dataset to ensure both authenticity and richness. Specifically, we collected 202,760 image-text pairs from 15 top-tier computer science conferences and refined them into 37,607 high-quality charts with contextual information. Building on this foundation, we introduce the Scientific Chart QA (SCI-CQA) benchmark, which includes a diverse evaluation set of 5,629 carefully constructed questions, a training set generated via a structured annotation pipeline, and a human-inspired evaluation framework. Through comprehensive experiments involving 13 open-source and 8 proprietary models, we uncover critical limitations in current model performance. A key finding is that fine-grained visual perception is essential for accurate chart comprehension. To address this, we propose the Visual Insight Extraction Module (VIEM), which leverages saliency cues-such as keypoints and OCR signals-to highlight essential visual details. Integrating VIEM into base models yields consistent improvements on fine-grained question answering tasks. Our code and dataset will be publicly released to support further research and promote progress in scientific chart understanding.