Preprint
Aug 2026
Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
This work introduces CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits, enabling us to evaluate and improve methods for explaining LLM behaviors.
Adam Karvonen, Euan Ong, Subhash Kantamneni et al.
· 2 citations