Skip to content

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Sep 2026 · 0 citations · 42 references
Computer Science

TL;DR

Findings reveal sustained, dependency-consistent evidence integration, rather than isolated fact retrieval, as a key bottleneck for current deep research agents.

Abstract

Deep research agents are increasingly capable of web search, tool use, multimodal evidence analysis, and information synthesis. However, existing benchmarks mainly evaluate medium-horizon exploration and rarely test whether agents can sustain long, dependency-heavy research processes. We introduce Mr. LHDR (Multimodal real-world Long-Horizon Deep Research), a benchmark for evaluating real-world deep research over long, irreducible chains of interdependent evidence across eight categories. Each question is constructed from a hidden Node-Relation graph and requires an average of 12.1 necessary intermediate conclusions with a mean dependency depth of 10.4 before reaching a short, unique, and verifiable answer. Questions incorporate multimodal evidence, including images, maps, PDFs, logos, charts, tables, and video frames, with at least one non-text element that changes the reasoning state. Mr. LHDR evaluates both final answers and the correctness of intermediate conclusions under annotated dependencies. We evaluate general models, deep research systems, and agent frameworks using Overall Accuracy (OA), Strict Accuracy (SA), Checklist Score (CS), and Dependency-Aware Checklist Score (DACS). Results show that even the strongest system achieves only 43.1% OA and 34.3% SA, indicating that final-answer accuracy substantially overestimates complete research success. Removing images reduces DACS by 12.6 points, demonstrating the importance of multimodal evidence, while SA consistently declines as reasoning chains become longer. These findings reveal sustained, dependency-consistent evidence integration, rather than isolated fact retrieval, as a key bottleneck for current deep research agents.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in video understanding, existing benchmarks remain confined to simple scene-...

Seoyeon An, Hyeonseo Jang, Minsu Kim et al. · 0 citations
#artificial intelligence Preprint Aug 2026

WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

WeAgent-Harness, a multimodal agentic harness that supports native text-vision interaction and runtime recovery, and WeAgent-MMSearch, an integrated system spanning data construction, agentic post-training, and multimodal rollout that outperform similarly sized open-source models and rival models with roughly ten times...

Zongkai Liu, Hui Zhang, Li-Qiang Niu et al. · 0 citations
#artificial intelligence Review Sep 2026

LongCat-DeepResearch Technical Report

We present LongCat-DeepResearch, a deep research system that combines an enhanced LongCat model with a multi-agent workflow for producing comprehensive, evidence-grounded reports. The workflow separates global planning from detailed investigation and coordinates revision at the section level. Multiple planning agents f...

Mei Zhu, Yue-Ya Xu, Wan-Li Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

WorldBench: Culturally Grounded Benchmark for Multilingual Agents

WorldBench is presented: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions, and Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and ot...

Leonardo Ranaldi, Sherrie Shen, Jushi Kai et al. · 0 citations
Preprint Sep 2026

AdaptArena: Evaluating Test-Time Personalization of Web Agents

Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings where user intentions are underspecified and preferences are heterogeneous. In practice, users rarely provide explicit profiles, requiring agents to infer latent prefere...

Dong-Chan Shin, Xing Han Lù, Jiaqi Deng et al. · 0 citations
#artificial intelligence Preprint Sep 2026

IterSynth: Rethinking Deep Search Agents via Role-Decoupled Iterative Synthesis

Deep search requires LLM agents to decompose complex queries, search for evidence, and synthesize grounded answers, yet existing ReAct-style agents suffer from two limitations: role coupling, where one policy must handle planning, evidence use, and synthesis; and context accumulation, where growing search histories int...

Xing-Yu Wu, Yuchen Yan, Zhengxi Lu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.