Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots
Abstract
Large language models (LLMs) frequently endorse and elaborate on users’ delusional beliefs, a failure mode termed psychogenicity in the Psychosis-Bench study of Au Yeung et al., whose framing we adopt. We present the first systematic evaluation of whether anti-sycophancy interventions transfer to psychosis-relevant contexts. Across 1,280 experiments spanning 10 conditions, 8 frontier LLMs, and 16 clinically derived Psychosis-Bench scenarios, a combined anti-sycophancy prompt reduces mean Delusion Confirmation Scores by 73.8% (paired \(t(127)=11.74\) , \(p<10^{-21}\) , Cohen's \(d=1.04\) ). Adding a domain-general self-reflection prompt yields a 77.0% reduction and raises Safety Intervention rates by 66.9%, delivered entirely as a system prompt. Classifier-based guardrails (Llama Guard 3) flag only 5 of 3,072 evaluated turns; a reasoning guardrail (o4-mini) flags \(14\times\) more. Ablations isolating either mechanism alone plateau at \(\approx 46\%\) reduction, establishing anti-sycophancy prompting as a necessary foundation that add-on mechanisms augment but cannot replace.