Epistemic State Representations in Large Language Models
Large language models often generate confident but fabricated content, yet whether they maintain internal representations of their own epistemic states is unknown. To address this question, contrastive activation vectors were extracted for 15 epistemic states from five language models, using 100 matched present-neutral vignette pairs per state. A low-dimensional geometry emerged in all five models: a dominant confidence axis groups states by certainty irrespective of category, and a secondary axis distinguishes self-directed from world-directed uncertainty (Fisher combined p = 0.011). Steering and targeted activation patching along these directions causally alter behavior, at times rescuing correct answers on previously confabulated questions. When combined with output-level confidence, the vectors yield a confabulation detector with an ROC-AUC of 0.89-0.92. We anticipate that these vectors will be useful for diagnosing and correcting epistemic failure in deployed language models.