Skip to content

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

Sep 2026 · 0 citations · 77 references
Computer Science

TL;DR

Sci-MMR is introduced, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions, and it is found that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents.

Abstract

Autonomous research agents are increasingly expected to search the literature, analyze experimental evidence, and generate scientific hypotheses. These capabilities require multi-step evidence grounded reasoning that progressively acquires, integrates, and verifies evidence before reaching a conclusion. Existing multimodal benchmarks, however, largely evaluate final-answer accuracy, leaving open whether predictions are actually supported by traceable scientific evidence. We introduce Sci-MMR, a benchmark for multi-step evidence-grounded scientific reasoning built on structured argument graphs linking scientific claims, citation-grounded knowledge, visual evidence, and supporting regions. Sci-MMR comprises 235 multi-hop reasoning tasks spanning four scientific disciplines, with an average of nine figure panels per task. Evaluating eight frontier multimodal models, we find that answer accuracy consistently exceeds complete-evidence recovery rate by more than 20%, revealing a substantial gap that answer-only evaluation is structurally unable to capture. Through controlled interventions, we identify two fundamental bottlenecks. First, evidence acquisition: models struggle to extract complete structured evidence from scientific figures, accounting for 57.2% of failures. While cropping tools yield modest gains (+4.5 points), providing gold evidence improves accuracy by up to 37.0 points, indicating difficulty in assembling complete multi-region evidence. Second, evidence integration: models struggle to translate available evidence into correct conclusions, accounting for 31.8% of failures, while even with gold evidence the strongest model achieves only 69.1% accuracy on the hardest tasks. These findings indicate that current answer-centric benchmarks substantially overestimate the evidence-grounded reasoning capabilities of multimodal research agents

View source

Similar papers

Preprint Aug 2026

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

This work introduces Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains, and establishes a rubric-based evaluation protocol, showing that advances in visual realism have not yet translated into reliable modeling of scientific and causal dyn...

Diandian Zhang, Tingyu Song, Linbo Fu et al. · 1 citation
Preprint Sep 2026

SciLENS: RL-Driven Autonomous Agents for Scientific Localized Evidence Navigation and Synthesis

Scientific literature synthesis agents increasingly rely on proprietary online services, limiting reproducibility, privacy, and offline deployment. To address this challenge, we introduce SciLENS Scientific Localized Evidence Navigation and Synthesis), a fully local autonomous agent framework operating on a dual-tier i...

Le-Qi Zheng, Jin-Bo Su, Yu-Ying Li et al. · 1 citation
#artificial intelligence Preprint Sep 2026

SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores

Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether their conclusions are scientifically supported. We formulate evidence-grounded multimodal scientific analysis, requiring agents to produce executable analyses and claims supp...

Bowen Liu, Shuo Nie, Bo-Dong Du et al. · 0 citations
#artificial intelligence Review Sep 2026

DISCERN: Can AI Agents Work Like Scientists and Guide Discovery?

Reliable automated research requires agents to vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence, potentially reducing routine scientific workload while allowing scientists to focus on interpretation and discovery. Existing benchmarks often only assess analytical task completion or hyp...

Nan Huang, Mario Tapia-Pacheco, Kun Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Marathon of Scientific Reasoning: Robustness of Scientific Agents to Perturbations in Multi-Turn Interactions

Large language model (LLM)-based scientific agents are increasingly used for scientific problem solving, yet their robustness to imperfections arising during multi-turn interactions remains poorly understood. We introduce \textsc{SciARP} (\textbf{Sci}entific \textbf{A}gent \textbf{R}obustness to \textbf{P}erturbations)...

Xiaoting Lyu, Xin-Bo Ma, Yu-Fei Han et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era

AgentIdeaBench is introduced, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration, and Scientific World Modeling is explored, a generation-time loop that refines a draft hypothesis through structured thought experiments.

Yunxiang Mo, Tianshi ZHENG, Yi-Sen Gao et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.