Aug 2026· Discover Artificial Intelligence· 0 citations
TL;DR
A unified symbolic, behavioral, and mechanistic framework that connects symbolic triggers with internal failure dynamics in transformer architectures and provides an interpretable basis for diagnosing and stabilizing symbolic reasoning in LLMs is introduced.
Abstract
Large Language Models (LLMs) tend to hallucinate when processing symbolically complex linguistic structures. Existing literature evaluates hallucination either through their mechanistic interpretability or at the behavioral output level, but hardly links the symbolic triggers to their layer-wise representational causes. This paper introduces a unified symbolic, behavioral, and mechanistic framework that connects symbolic triggers with internal failure dynamics in transformer architectures. The study evaluates five open-weight LLMs across QA, MCQ, and Odd-One-Out formats on the HaluEval and TruthfulQA datasets, focusing on negation, exceptions, modifiers, numbers, and named entity cues. The results show that hallucination rates remain high across all model scales, with all symbolic categories exhibiting high hallucination rates, and exceptions and numbers often showing comparable or higher values across models. Constrained task formats reduce surface errors but preserve failure patterns, indicating representational instability rather than purely decoding artifacts. Layer-wise analysis shows peak symbolic attention variance in early transformer layers (2–4), after which these patterns persist across deep layers. The consistency of this behavior across architectures suggests that hallucination is strongly associated with weakness in symbolic encoding. The framework provides an interpretable basis for diagnosing and stabilizing symbolic reasoning in LLMs.
A lightweight linear detector is built on top of Role-Break that requires no fine-tuning of the VLM, whose feature dimension stays below 5,000 and reaches an average AUROC of 93.23 across six VLMs and four benchmarks.
Mingyu Wang, Weilin Jin, Wenbo Li et al.· 0 citations
ReWEIGH is a training-free decoding intervention that aggregates vocabulary ranks across visual positions and compares each candidate with a token-specific reference estimated from unlabeled images and applies a bounded penalty only to candidates that fall below their reference.
This review provides systematic theoretical support for industrial RAG model selection and optimization and summarizes existing research gaps, including lightweight deployment and multimodal expansion, and proposes future research directions for trustworthy RAG systems.
Shujing Liu· Applied and Computational En...· 0 citations
The D-Score is introduced, a simple spectral statistic computed from a single forward pass that is used as a hallucination score, classifying an input text as hallucinated when its D-Score is larger than a pre-defined quantity.
Bianca Raimondi, Davide Evangelista, Maurizio Gabbrielli et al.· 0 citations
This review paper provides a comprehensive overview of hallucinations in GAI and LLMs, and synthesizes a range of correction and mitigation techniques, from proactive measures during training to hybrid approaches that combine detection and intervention.
M. Naser· Language Resources and Evalu...· 0 citations
Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model predictions are not grounded in visual evidence, so a fully black-box framework that models hallucination as a structured uncertainty pattern is proposed.
Amir Mohammad Ezzati, Kiyan Rezaee, Bardiya Kariminia et al.· 0 citations