The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models
Bounded margins mitigate confident hallucinations during post-training, implemented through an entropy-dependent margin bound in direct preference optimization (DPO) and shown to mitigate confident hallucinations during post-training.