This work builds a trajectory-level diagnostic pipeline that evaluates both answer correctness and evidence-grounding quality under a unified ReAct-style scaffold and introduces a six-category taxonomy covering modality shortcuts, phantom grounding, wrong-evidence-right-answer cases, over-retrieval laundering, cross-modal contradiction, and provenance hallucination.
Video understanding is increasingly performed by multi-stage LLM agents that separate temporal grounding, visual observation, and reasoning. Yet these stages are typically evaluated on different benchmarks and distributions, making it difficult to determine where hallucinations originate. We first organize existing ben...
responses depending on domain, retriever quality, and
model family. This paper reviews the literature on why RAG systems continue to hallucinate even when correct evidence is
available in context, organizes the reported causes into a five-part taxonomy (retrieval failure, conflicting evidence,
unfaithful generation, ov...
Sanchita H., Skandamahima V. M., Akshitha Katkeri· International Journal of Inn...· 0 citations
When strong multimodal models are widely available, progress requires new scientific methodologies beyond benchmark scores---using models as instruments for understanding behavior. We address this by asking: can we use large vision-language models (LVLMs) as experimental instruments for studying their own failure dynam...
Zhi-Peng Zhao, Wen-Xu Wang, Peishun Liu et al.· 0 citations
Savor is introduced, a training framework that augments the output schema with token and answer confidence, optimises the policy with a Group Relative Policy Optimisation objective that penalises calibration error and poor abstention decisions, and uses the learned confidence at inference time to revisit visual evidenc...
Zian Ding, Zi-Lin Zhao, Ying-Jie He et al.· 0 citations
Multimodal large language models generate natural-language responses from visual inputs, yet may mention objects absent from an image. In medication assistance, accessible perception, and environmental decision-making, such hallucinations can create real-world safety risks. We propose Semantic-Spatial Agreement Verific...
Zi-Heng Ren, Qian Gao, Jun Fan et al.· 0 citations
The results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.
Heng-Yuan Xu, Wei Cheng, Yu-Meng Ji et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.