Skip to content

The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models

Sep 2026 · 0 citations · 49 references
Computer Science

TL;DR

Bounded margins mitigate confident hallucinations during post-training, implemented through an entropy-dependent margin bound in direct preference optimization (DPO) and shown to mitigate confident hallucinations during post-training.

Abstract

Large language models (LLMs) can produce factually incorrect answers with high confidence, undermining their reliability and limiting the effectiveness of uncertainty-based error detection. While prior research attributes confident hallucinations to factors such as missing knowledge in training data, reasoning errors, or stochastic decoding, we uncover that post-training alignment itself is a primary driver of these errors, a phenomenon we call the \textbf{Alignment Paradox}. Across five model families evaluated on factual benchmarks, unaligned base models produce few high-confidence errors on long-tail factual queries, whereas instruction-tuned models multiply high-confidence errors ($p \ge 0.95$) by more than an order of magnitude (10$\times$ to 35$\times$). Layer-wise probing with the Logit Lens reveals that this overconfidence emerges in late layers, where wrong-answer margins expand past 4.0 points after remaining near zero across early and intermediate layers. These findings motivate limiting margin growth during post-training. We implement this principle through an entropy-dependent margin bound in direct preference optimization (DPO). In multi-epoch experiments with Mistral-7B, the bounded objective reduces high-confidence errors by up to 35.3\% relative to standard DPO while maintaining performance on evaluated general reasoning benchmarks. These results show that bounded margins mitigate confident hallucinations during post-training.

View source

Similar papers

#machine learning Preprint Sep 2026

Attention Dispersion as a Diagnostic Signal for Hallucination in Large Language Models

Large Language Models (LLMs) frequently exhibit hallucinations, presenting a major barrier to reliability in complex reasoning tasks. While traditional detection methods rely on output-based confidence metrics, these logits are often miscalibrated by modern alignment techniques. In this paper, we investigate the tempor...

S. More, Tanuja S. Pawar · 0 citations
#artificial intelligence Preprint Sep 2026

Domain-Specific Hallucination Detection in Large Language Models

Large language models generate fluent text that can contain unfaithful claims -- a phenomenon known as hallucination. We present a multi-signal detection pipeline combining fine-tuned DeBERTa-v3 classification, Monte Carlo (MC) Dropout uncertainty quantification, and temperature-scaled calibration for response-level ha...

Varun Teja Chundru, Debasmita Biswas · 0 citations
#artificial intelligence Preprint Sep 2026

The Detectability Gap: Hidden Heterogeneity in Hallucination Detection Across Language Models

Sampling based consistency is widely used for hallucination detection, yet aggregate performance can conceal systematic differences in which errors are detectable. This work studies that heterogeneity across four language models and three factual question answering datasets. Partitioning hallucinations by answer agreem...

Pranav Darshan, A. Pranav, T. SravanKarthick et al. · 0 citations
Open access Sep 2026

Large Language Models Create Hallucinations in Response to Negated Text

Large language models (LLMs) have achieved significant advancements in natural language processing tasks, but they remain prone to generating hallucinations—outputs that are logically inconsistent or factually incorrect. While previous research has primarily focused on hallucinations in affirmative contexts, how negate...

Jaehyung Seo, Hyeonseok Moon, Heu-Jeoung Lim · 0 citations
Conference 2026

ContrastSFT: Contrastive Logit Regularization Supervised Fine-Tuning for Mitigating Hallucinations in Large Language Models

Large Language Models (LLMs) have achieved remarkable success in natural language generation but remain prone to hallucinations—generating content that is fluent but factually incorrect. While recent inference-time interventions like Contrastive Decoding (CD) effectively mitigate this by penalizing tokens favored by...

Si-Fan Zhou · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.