Bounded margins mitigate confident hallucinations during post-training, implemented through an entropy-dependent margin bound in direct preference optimization (DPO) and shown to mitigate confident hallucinations during post-training.
Abstract
Large language models (LLMs) can produce factually incorrect answers with high confidence, undermining their reliability and limiting the effectiveness of uncertainty-based error detection. While prior research attributes confident hallucinations to factors such as missing knowledge in training data, reasoning errors, or stochastic decoding, we uncover that post-training alignment itself is a primary driver of these errors, a phenomenon we call the \textbf{Alignment Paradox}. Across five model families evaluated on factual benchmarks, unaligned base models produce few high-confidence errors on long-tail factual queries, whereas instruction-tuned models multiply high-confidence errors ($p \ge 0.95$) by more than an order of magnitude (10$\times$ to 35$\times$). Layer-wise probing with the Logit Lens reveals that this overconfidence emerges in late layers, where wrong-answer margins expand past 4.0 points after remaining near zero across early and intermediate layers. These findings motivate limiting margin growth during post-training. We implement this principle through an entropy-dependent margin bound in direct preference optimization (DPO). In multi-epoch experiments with Mistral-7B, the bounded objective reduces high-confidence errors by up to 35.3\% relative to standard DPO while maintaining performance on evaluated general reasoning benchmarks. These results show that bounded margins mitigate confident hallucinations during post-training.
Large Language Models (LLMs) frequently exhibit hallucinations, presenting a major barrier to reliability in complex reasoning tasks. While traditional detection methods rely on output-based confidence metrics, these logits are often miscalibrated by modern alignment techniques. In this paper, we investigate the tempor...
This work introduces a two-stage keyword-perturbation method for hallucination detection and extends the same probabilistic framework to four hallucination regimes: knowledge deficit, wrong knowledge, context distraction, and unstable inference.
Xu-Han Tong, Hao-Yue Bai, Da-Wei Zhou et al.· 0 citations
Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level ha...
Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreem...
Pranav Darshan, A. Pranav, T. SravanKarthick et al.· 0 citations
Large language models (LLMs) have achieved significant advancements in natural language processing tasks, but they remain prone to generating hallucinations—outputs that are logically inconsistent or factually incorrect. While previous research has primarily focused on hallucinations in affirmative contexts, how negate...
Jaehyung Seo, Hyeonseok Moon, Heu-Jeoung Lim· ACM Transactions on Knowledg...· 0 citations
Large Language Models (LLMs) have achieved remarkable success in natural language generation but remain prone to hallucinations—generating content that is fluent but factually incorrect. While recent inference-time interventions like Contrastive Decoding (CD) effectively mitigate this by penalizing tokens favored by...
Si-Fan Zhou· Poster Volume 0007 The 2026...· 0 citations