The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models
Bounded margins mitigate confident hallucinations during post-training, implemented through an entropy-dependent margin bound in direct preference optimization (DPO) and shown to mitigate confident hallucinations during post-training.
Qing-Jia Huang, Ya-Kai Li, Jian-Guo Wu et al.
· 0 citations