Large language models (LLMs) frequently produce source code that seems correct and well-formed, yet includes hallucinated elements that cause downstream test failures. In this study, we benchmark state-of-the-art uncertainty quantification methods and existing base-lines for the task of hallucination detection in source code and introduce a diff-based pipeline to construct a code dataset annotated with line-level hallucinations. Building on this, we train a lightweight Transformer-based detector that uses LLM internal representations to identify hallucinations, substantially outperforming existing methods across several code generation domains. The detector also shows particular promise for enabling self-correction in LLM-based coding agents. We release the first publicly available dataset of line-level code hallucinations, along with the corresponding source code and trained hallucination detectors https://github.com/ datapaf/CodeHallucinationDetection
G. Andriushchenko, Roman Garaev, L. Rvanova et al.· Annual Meeting of the Associ...· 0 citations
An extensive experimental study is presented demonstrating that ARMT-augmented models process inputs well beyond their original context limits without degrading performance relative to in-limit baselines and need 30% less FLOPs while preserving baseline performance within the original context window.
Gleb Kuzmin, I. Rodkin, A. Bulatov et al.· 0 citations