Skip to content

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

Jul 2026 · arXiv.org · Vol abs/2607.20891 · 4 citations · 46 references
Computer Science

TL;DR

MisKnow-Agent is introduced, a controlled evaluation framework that constructs task-specific documents supporting manually audited false conclusions with controlled authority cues and source styles that evaluates DeerFlow and WebThinker with three backbone LLMs using a report-level false-conclusion adoption rate that counts only reports endorsing the false conclusion.

Abstract

Deep Research agents conduct long-horizon investigations by iteratively planning, retrieving evidence, and generating reports. However, it remains unclear whether they can resist apparently credible but factually false information introduced into these workflows. To study this failure mode, we introduce MisKnow-Agent, a controlled evaluation framework that constructs task-specific documents supporting manually audited false conclusions with controlled authority cues and source styles. Applied to the tasks from DeepResearch Bench, it generates 5,933 misleading documents after filtering. We evaluate DeerFlow and WebThinker with three backbone LLMs, together with Gemini Deep Research, using a report-level false-conclusion adoption rate (FCAR) that counts only reports endorsing the false conclusion. Across the configurations, introducing one misleading document increases the mean FCAR from 0\% in the no-injection control to 54.7\%. FCAR varies substantially with lifecycle stage and framework design, and also with source authority and presentation style, whereas search-result rank and additional documents beyond the first have limited influence. Although cross-model verification consistently classifies retained instances as misleading, Deep Research agents can still adopt the corresponding false conclusions during long-horizon research. Pre- and post-research defenses reduce FCAR but do not eliminate adoption, motivating continuous verification when evidence enters intermediate research states and final synthesis. To facilitate reproducibility, our code and dataset are publicly available at https://github.com/whfeLingYu/MisKnow-Agent and https://huggingface.co/datasets/whfeLingYu/Misleading_Knowledge, respectively.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

RINI: Seeing the Prior Is Not Enough

A research proposal can describe an established mechanism correctly while claiming to introduce it. We study whether providing the earlier paper corrects such contribution claims. Three controlled experiments compare proposals generated with a contribution-bearing prior and a same-topic control. Providing the prior yie...

Hong-Yi Du, Tian-Yi Zhang, Heng Wang et al. · 0 citations
Preprint Aug 2026

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

A probe corpus of 42 retracted, fraudulent, and pseudoscientific papers is paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing, indicating an urgent need for guardrail infrastructure for scientific deployment of language models.

V. Rodionov, Shamil Assylbekov · 0 citations
#artificial intelligence Review Oct 2026

Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents

Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three...

Jun-Yu Guo, Shangding Gu, Ming Jin et al. · 0 citations
Preprint Aug 2026

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

The Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable.

Zi-Yue Wang, Aomufei Yuan, Yi-Ran Yao et al. · 0 citations
Preprint Aug 2026

Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research

An evaluation method is proposed that pinpoints which agent introduced each error by locally testing agent invocations for faithfulness and verifiability relative to their own inputs and proposes a four-type taxonomy to categorize the discovered errors: hallucination, uncited input reliance, uncited output, or insuffic...

Eran Hirsch, David Wan, Han Wang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.