Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations
Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what"truth"means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas whose beliefs clearly con...