Skip to content

MISHAP-Bench: A Hallucination Benchmark for Large Audio-Language Models

Sep 2026 · 0 citations · 66 references
Computer Science

TL;DR

This work introduces MISHAP-Bench, a comprehensive benchmark with 12,000 challenging open-ended question-audio pairs and a rigorous evaluation pipeline covering two hallucination categories, and proposes a groundedness judge that uses reference rubrics and judge prompts guided by human annotations.

Abstract

Large audio-language models (LALMs) produce fluent responses about audio but often hallucinate by making plausible yet ungrounded claims. Existing audio hallucination benchmarks mainly measure response correctness, leaving it unclear whether an LALM hallucinates or simply fails to understand the audio. We challenge correctness-based evaluation by defining two hallucination categories: (i) context, where claims are not grounded in the audio; and (ii) knowledge, where claims about audio-related topics lack support from externally verifiable facts. We introduce MISHAP-Bench, a comprehensive benchmark with 12,000 challenging open-ended question-audio pairs and a rigorous evaluation pipeline covering both categories. To evaluate open-ended responses, we propose a groundedness judge that uses reference rubrics and judge prompts guided by human annotations. We extensively evaluate ten state-of-the-art LALMs and show that hallucination remains substantial. Even a frontier model such as Gemini 3.7 Flash reaches a hallucination rate of 36.5%. We further adapt and benchmark four mitigation methods from multiple domains for LALMs. Despite some improvements, effective hallucination mitigation remains an open challenge. Finally, we call on the community to evaluate hallucination and benchmark mitigation methods with MISHAP-Bench.

View source

Similar papers

Book Open access Aug 2026

Who's Adam? Benchmarking Hallucinations in Scientific Dialogue

ADAM-Bench (Auditing Dialogue Assertions with Multimodal Evidence), a benchmark for paper-grounded hallucinations in scientific dialogue, is introduced and two tasks are defined: hallucination detection and minimal evidence set localization.

Ze-Xing Zhang, Tian-Yang Lei, Ke-Wei Yang et al. · 0 citations
Open access Sep 2026

Large Language Models Create Hallucinations in Response to Negated Text

Large language models (LLMs) have achieved significant advancements in natural language processing tasks, but they remain prone to generating hallucinations—outputs that are logically inconsistent or factually incorrect. While previous research has primarily focused on hallucinations in affirmative contexts, how negate...

Jaehyung Seo, Hyeonseok Moon, Heu-Jeoung Lim · 0 citations
Preprint Sep 2026

Beyond Accuracy: ARIA-Rubrics for Evaluating Audio Reasoning in Large Audio Language Models

Large Audio Language Models (LALMs) have shown strong performance on audio reasoning benchmarks, but accuracy alone cannot distinguish true reasoning from superficial pattern matching, often overestimating reasoning ability since high scores may result from guessing rather than genuine audio understanding. Evaluating t...

Yu-Pei Li, Qi-Yang Sun, Mohamed Mady et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Domain-Specific Hallucination Detection in Large Language Models

Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level ha...

Varun Teja Chundru, Debasmita Biswas · 0 citations
Preprint Aug 2026

Hallucination Span Detection with Input-Side Evidence Alignment

This work introduces the task of hallucination span detection with input-side evidence alignment, which jointly identifies hallucinated spans and aligns output tokens with the corresponding input evidence.

Miyu Yamada, Yuki Arase · 0 citations
#computer vision Preprint Sep 2026

TAD: Token-Adaptive Contrastive Decoding with Confidence-Guided Gating for Hallucination Mitigation in Large Audio-Language Models

Large audio-language models (LALMs) can hallucinate audio objects, answering"yes"to absent sound events, thus undermining reliability in audio question answering. We propose Token-Adaptive Decoding (TAD), a training-free strategy for hallucination mitigation that grounds the initial yes/no decision by contrasting logit...

He-Yu Chang, Nianwen Si, Hao Zhang et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.