Skip to content
Conference Open access

SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration

Jul 2026 · Annual Meeting of the Association for Computational Linguistics · pp. 22249-22273 · 1 citation · 28 references
Computer Science

TL;DR

SciExplore is introduced, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks.

Abstract

Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks primarily emphasize general-domain retrieval or static scientific question answering, and therefore fail to assess key capabilities required in realistic scientific research workflows. We introduce SciExplore, a benchmark designed to evaluate scientific information-seeking and reasoning capabilities of LLMs and agents. SciExplore comprises four task types covering 103 expert-curated tasks across more than ten scientific disciplines: scientific database navigation, ambiguous literature retrieval, missing reference completion, and cross-source structured knowledge synthesis, which probe progressively higher-level abilities from entity-level reasoning and document-level identification to evidence-level grounding and domain-level synthesis. We evaluate over ten state-of-the-art LLMs and autonomous agents on SciExplore, revealing substantial performance gaps with performance degrading sharply as task complexity increases and extremely low accuracy on the most challenging structured synthesis tasks. These results highlight significant limitations of current models and agents in realistic scientific information-seeking scenarios.

Read PDF

Similar papers

Preprint Jul 2026

SciDataSailor: Deep Scientific Data Exploring

This work presents SciDataSailor, a framework for synthesizing tool-interactive trajectories by balancing broad exploration with targeted exploitation and presents SciDataSailor, a framework for synthesizing tool-interactive trajectories as Monte Carlo Tree Search (MCTS) with four task-specific mechanisms.

J. Rao, Yicheng Qiu, Chi Zhang et al. · 1 citation
Book Open access Jul 2026

HiRA: Decoupling Planning and Execution with Hierarchical Reasoning in Deep Search

Experiments show that HiRA significantly outperforms state-of-the-art RAG and agent-based systems, highlighting the effectiveness of decoupled planning and execution for multi-step information seeking tasks.

Jiajie Jin, Xiaoxi Li, Yuyao Zhang et al. · 0 citations
#artificial intelligence Review Jul 2026

ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System

ReasFlow is introduced, an end-to-end autonomous agent system for reasoning-centric scientific discovery that operationalizes a collaborative paradigm where the human expert acts as Principal Investigator while the agent executes rigorous derivations as a capable graduate student.

Yutong He, Daibo Li, Guohong Li et al. · 1 citation
Preprint Jul 2026

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their navigation and task completion performance has emerged as a critical research priority. However, existing benchmarks exhibit fundamental limitations. First, they suffer from insufficient scale and limited domain diversity, constraining comprehensive evaluation of cross-domain generalization. Second, prevailing LLM-as-Judge evaluation methodologies inadequately capture fine-grained interaction semantics, particularly regarding precise query formulation and filtering operations. Third, current benchmarks predominantly emphasize navigation success metrics while neglecting critical requirements for real-world deployment scenarios. To address these limitations, we introduce WebRetriever, a large-scale benchmark encompassing 800 websites and 1,550 tasks across diverse domains, including consumer, professional, and enterprise sectors, with comprehensive coverage of user intent patterns. We propose NavEval (Navigation Evaluation), a novel LLM-as-Judge framework that leverages rich interaction context beyond visual screenshots, achieving state-of-the-art alignment with human judgment across multiple evaluation datasets. Furthermore, we establish three complementary evaluation protocols that collectively provide holistic assessment of web agent capabilities: navigation proficiency, knowledge-assisted interaction, and end-to-end task completion with information extraction. Extensive experimental analysis reveals substantial performance disparities across evaluation protocols, demonstrating that navigation success alone is an insufficient predictor of real-world application effectiveness. WebRetriever delivers fine-grained diagnostic insights into agent capabilities and establishes a rigorous foundation for advancing web agent research and development.

Wei Dong, Tianyu Fu, Zhe Yu et al. · 0 citations
Review Open access 2024

Agentic AI Framework for Autonomous Scientific Research Assistance

Experimental results indicate that coordinated autonomous agents significantly reduce research time, improve workflow consistency, enhance knowledge discovery, and increase scientific productivity compared with conventional AI-based research assistants.

Anatoly Kitov, M. Kartsev · 0 citations
Preprint Jul 2026

SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration

This work introduces SearchOS, a system-level multi-agent framework that turns fragile, implicit search progress into explicit, persistent, and shared state, and introduces a Search Tool Middleware Harness that intercepts model and tool interactions to record grounded evidence and react to stalls or budget exhaustion.

Yuyao Zhang, Junjie Gao, Zhengxian Wu et al. · 0 citations