Skip to content

Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation

2026 · arXiv.org · Vol abs/2607.19793 · 0 citations · 9 references
Computer Science

TL;DR

This work builds a trajectory-level diagnostic pipeline that evaluates both answer correctness and evidence-grounding quality under a unified ReAct-style scaffold and introduces a six-category taxonomy covering modality shortcuts, phantom grounding, wrong-evidence-right-answer cases, over-retrieval laundering, cross-modal contradiction, and provenance hallucination.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models

Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to determine where hallucinations originate. We first organize existing ben...

Shu-Zhi Gong, F. Sun, Yuansan Liu · 0 citations
Review Open access Sep 2026

Why Retrieval Doesn't Cure Everything: A Review of Hallucination in Retrieval-Augmented Generation

responses depending on domain, retriever quality, and model family. This paper reviews the literature on why RAG systems continue to hallucinate even when correct evidence is available in context, organizes the reported causes into a five-part taxonomy (retrieval failure, conflicting evidence, unfaithful generation, ov...

Sanchita H., Skandamahima V. M., Akshitha Katkeri · 0 citations
Preprint Sep 2026

What Do Hallucinations Reveal About Multimodal Reasoning? Diagnosing Visual Grounding Failures via Contrastive Decoding Probes

When strong multimodal models are widely available, progress requires new scientific methodologies beyond benchmark scores---using models as instruments for understanding behavior. We address this by asking: can we use large vision-language models (LVLMs) as experimental instruments for studying their own failure dynam...

Zhi-Peng Zhao, Wen-Xu Wang, Peishun Liu et al. · 0 citations
Preprint Sep 2026

SAVOR: Self-Aware Visual Grounding via Confidence-Calibrated Reinforcement Learning for Multimodal Hallucination Mitigation

Savor is introduced, a training framework that augments the output schema with token and answer confidence, optimises the policy with a Group Relative Policy Optimisation objective that penalises calibration error and poor abstention decisions, and uses the learned confidence at inference time to revisit visual evidenc...

Zian Ding, Zi-Lin Zhao, Ying-Jie He et al. · 0 citations
Preprint Sep 2026

Semantic-Spatial Agreement Verification for Mitigating Object Hallucination in Multimodal Large Language Models

Multimodal large language models generate natural-language responses from visual inputs, yet may mention objects absent from an image. In medication assistance, accessible perception, and environmental decision-making, such hallucinations can create real-world safety risks. We propose Semantic-Spatial Agreement Verific...

Zi-Heng Ren, Qian Gao, Jun Fan et al. · 0 citations
Preprint Aug 2026

Aphanta: Diagnosing Task-Aligned Image-Edited Intermediates for Multimodal Reasoning

The results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.

Heng-Yuan Xu, Wei Cheng, Yu-Meng Ji et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.