Skip to content
Book Open access

SciChart: Visual Question Answering and Reasoning for Scientific Spectral Chart

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 12751-12760 · 1 citation · 51 references

TL;DR

This work designs two tasks, basic question answering (BasicQA) and reasoning-based question answering (ReaQA), to evaluate the models' ability to directly extract information from charts, and understand the textual and visual information for reasoning.

Abstract

Charts play a key role in scientific research, offering a concise and visual way to present complex data. For Multimodal Large Language Models (MLLMs), the ability to comprehend charts is critical, as it requires both visual perception and reasoning that bridges graphical and textual information. However, existing chart question answering (QA) datasets are monolingual with simple questions, making current evaluation benchmarks inadequate for the rapid advancements in MLLM performance. Therefore, we propose a multilingual scientific spectral Chart QA dataset, termed SciChart. We design two tasks, basic question answering (BasicQA) and reasoning-based question answering (ReaQA), to evaluate the models' ability to 1) directly extract information from charts, and 2) understand the textual and visual information for reasoning. We build 1,100 ReaQA and over 10,000 BasicQA samples. All samples are manually curated and annotated by human experts. We also conduct extensive experiments with state-of-the-art models to establish SciChart benchmarks. Experimental results show a huge gap between the performance of existing models (Claude-3.7 45.12%) and humans (83.84%).

Read PDF

Similar papers

Sep 2026

Advancing Scientific Chart Understanding: The SCI-CQA Benchmark and Beyond.

In real-world applications, current multimodal large models are often overestimated in their ability to understand scientific charts. To assess their true capabilities and identify key performance bottlenecks, we conducted an in-depth study on scientific chart understanding. Charts in scientific literature often featur...

Ling-Dong Shen, Qigqi, Kun Ding et al. · 0 citations
#graph neural networks Review Open access Sep 2026

Exploring multi-answer visual question answering with object detection: a systematic review

This systematic literature review (SLR) focuses on multi answer VQA systems and the use of object detection, following the PRISMA 2020 guidelines and proposes a taxonomy of multi-answer VQA organized along four dimensions.

Nida Hasanati, Taufik Djatna, Imas Sukaesih Sitanggang et al. · 0 citations
Conference Open access Aug 2026

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951...

Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade et al. · 1 citation
Book Open access Oct 2026

MapQA: A Map-Question-Answering Benchmark for Visual Language Model Reasoning

Maps are central to how humans make sense of the world, from navigation and environmental monitoring to military planning and historical interpretation. Yet despite rapid progress in large multimodal models (LMMs), these systems continue to struggle with interpreting maps – an essential skill for visual reasoning that...

Christian M. Arnold, Andrew Alini, Abdulrahman Alabdulkareem et al. · 1 citation
#artificial intelligence Preprint Sep 2026

DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents

This work introduces DocHop, a benchmark for integrated chart--context reasoning in document-style images and constructs DocHop via a stochastic logic-first generation pipeline with controllable reasoning depth and visual density, to enable systematic evaluation.

Zhuoran Yu, Le Thien Phuc Nguyen, Jaden Park et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.