Skip to content

SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs

Sep 2026 · 0 citations · 38 references
Computer Science

TL;DR

This work introduces SearchAtlas, a framework that converts search trajectories into structured graphs whose edges represent how evidence is propagated across the reasoning trace, from the query that retrieves it to the final answer.

Abstract

LLM search agents are often evaluated on final-answer accuracy, overlooking the process. Analyzing a search strategy requires understanding how credible evidence is retrieved to address question constraints. This valuable information is buried in raw search trajectories that are long and difficult to parse. We introduce SearchAtlas, a framework that converts search trajectories into structured graphs whose edges represent how evidence is propagated across the reasoning trace, from the query that retrieves it to the final answer. Our automated parsing pipeline achieves a mean edge F1 of 86.0% against human-annotated graphs and remains consistent across repeated runs. We analyze five search agents on three benchmarks, revealing systematic differences in search scale and evidence aggregation. SearchAtlas exposes fragmented answer support, question constraints that do not reach the answer, and unverified parametric knowledge entering the response. These process failures are strongly associated with incorrect answers, even more so than an LLM judge given either the raw trajectory or the ordered query list, suggesting that the constructed graphs provide useful interpretability. Moreover, an audit of cases in which process-diagnostic scores disagree with final-answer correctness shows that they capture information not reducible to answer accuracy.

View source

Similar papers

Preprint Aug 2026

EviGraph: Towards Verifiable Evidence Construction for Information-Seeking Agents

EviGraph is presented, a deep-search framework that separates search execution from evidence recording while using a shared policy for the trainable roles, enabling reinforcement learning to directly supervise evidence construction rather than only the final answer.

Jia-Shun Chen, Yirong Mao, Wen-Hui Que · 0 citations
#artificial intelligence Preprint Sep 2026

Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses

Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled...

Mahsa Amani, Seungeon Lee, A. Dash et al. · 0 citations
Preprint Aug 2026

Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory

Experiments establish ReTree as an effective self-correcting memory abstraction for long-horizon search, and show that ReTree consistently outperforms Full-Trajectory ReAct in question-answering and search benchmarks.

Aijun Yang, Qianxue Guo, Ziyi Huang et al. · 0 citations
#natural language process... Preprint Sep 2026

EvoDuet: Bilevel Co-Evolution of Web Searching and Task Solving for Scientific Discovery

Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method that co-evolves sol...

Young-Jun Lee, Jinheon Baek, Soyeong Jeong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge

This work compares country-continent questions with noun, adjective, and code answers while keeping several fitted measurements distinct across Qwen, Llama, and Gemma to separate early readability, natural strength, causal steering, and later content dependence.

Wen-Lin Wei, Yuan Fang, Ren-He Jiang et al. · 0 citations
Preprint Aug 2026

EviReform: Evidence-Guided Query Reformulation for Multi-Hop Graph Retrieval

Results show that observed evidence can guide graph retrieval toward the part of a supporting chain left underspecified by the original question, and introduce EviReform, which separates revising the retrieval request from aggregating evidence in the graph.

Xinlong Xu, Yoshua Y. Li · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.