Skip to content

Look Before You Leap: Factual Decoding with Internal Attribution Signals

Sep 2026 · 0 citations · 41 references
Computer Science

TL;DR

This work identifies a factual-salient layer span within LLMs whose derived signal is selectively elevated for factual tokens and exhibits anomalous spikes at hallucination-prone steps, and proposes DescaPE, a decoding framework that leverages internal model signals to suppress hallucination-prone trajectories at inference time.

Abstract

Hallucination remains a critical challenge in large language models (LLMs), where early factual errors compound through autoregressive generation in a snowballing effect that neither post-hoc correction nor weight-level intervention can effectively preempt. We propose DescaPE (DEcoding Signal Control Against Path Error-snowballing), a decoding framework that leverages internal model signals to suppress hallucination-prone trajectories at inference time. Through sliding-window MLP ablation, we identify a factual-salient layer span within LLMs whose derived signal is selectively elevated for factual tokens and exhibits anomalous spikes at hallucination-prone steps. We train a lightweight probe to approximate this signal from a single forward pass and integrate it into candidate scoring to penalize high-risk continuations while rewarding factually grounded ones. Experiments across five factuality benchmarks on three LLMs demonstrate that DescaPE achieves factuality improvements over decoding-time baselines in multiple settings, while incurring only 1.10x latency overhead in our efficiency evaluation. Our code is available at https://github.com/hayeonggg/DESCAPE.

View source

Similar papers

Open access Aug 2026

Towards Trustworthy Large Language Models

An integrated conceptual frame-work that couples attention- and perturbation-based explainability with lightweight hallucination-detection signals and token-efficient inference strategies is presented, and a set of cross-cutting consistency metrics are instrumented with a set of cross-cutting consistency metrics.

Sakshi Parate, Shreyans Sanyal · 0 citations
#artificial intelligence Preprint Oct 2026

External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing

As Large Language Models (LLMs) increasingly serve as foundational reasoning engines, their tendency to hallucinate remains a critical vulnerability. While recent internal state probes offer a promising alternative to slow external retrieval systems, they largely reduce hallucination detection to a token-wise binary cl...

Kingshuk Gupta, Davide Buscaldi · 0 citations
#natural language process... Preprint Sep 2026

SFAD: Speculative Factuality-Aware Decoding

SFAD is presented, a speculative decoding framework that enhances contextual faithfulness without inference degradation and substantially improves faithfulness while achieving $2.48\times$ speedup, offering a practical solution for efficient LLMs.

Guan-Qiao Chen, Di Wang, Lijie Hu · 0 citations
Preprint Aug 2026

Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique

The Latent Critic is introduced, a lightweight low-rank adapter that operates concurrently with a frozen base LLM's generation to actively restructure the transformer's residual stream---amplifying latent grounding signals and translating them into localized, natural language feedback within a single sequence.

S. Vijayvargiya, R. Lokesh · 2 citations
#artificial intelligence Preprint Sep 2026

The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining

Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to answer reliably. A large body of existing work mitigates hallucination through detection or abstention mechanisms, but leaves open how models internally arrive at the decis...

Vy Nguyen, Zi-Qi Xu, Jeffrey Chan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

The Alignment Paradox: How Post-Training Amplifies Confident Hallucinations in Language Models

Bounded margins mitigate confident hallucinations during post-training, implemented through an entropy-dependent margin bound in direct preference optimization (DPO) and shown to mitigate confident hallucinations during post-training.

Qing-Jia Huang, Ya-Kai Li, Jian-Guo Wu et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.