Skip to content
Preprint

Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments

Aug 2026 · 2 citations · 31 references
Computer Science

TL;DR

This work introduces CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits, enabling us to evaluate and improve methods for explaining LLM behaviors.

Abstract

Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a"good"explanation? In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs. To this end, we introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits. This yields thousands of high-quality explanations for naturally-occurring model behaviors along with supporting counterfactual evidence. We apply CHIVE in two ways. First, we evaluate whether common LLM interpretability techniques improve an agent's ability to predict counterfactual model behaviors. Surprisingly, we find no uplift from any of the interpretability techniques studied. Second, we use CHIVE to generate training data. We find that training models to predict outcomes of CHIVE-generated counterfactual experiments generalizes to various out-of-distribution settings. Overall, CHIVE automatically discovers explanations of naturally-occurring LLM behaviors, enabling us to evaluate and improve methods for explaining LLM behaviors.

View source

Similar papers

#natural language process... Preprint Sep 2026

An Empirical Study of Counterfactual Self-Explanations in LLMs

This work evaluates ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales and shows that model scale is the strongest determinant of explanation quality.

Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis-Mastromichalakis et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Taking the Whys Seriously: Limitations of Counterfactual Explanations in Justification and Recourse

It is found that an organization's choices on measurement models for feature and labels, business requirements, model validation, and the metric of model success have as much or more impact on the generated counterfactuals as the specifics of the generating method.

Mattia Cerrato, Otto Sahlgren, Xenia Heilmann · 0 citations
#artificial intelligence Preprint Aug 2026

The Illusion of $\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs

A diagnostic benchmark for open-domain, open-form, long-horizon counterfactual causal reasoning, containing 220 what-if questions across STEM, HSS, and Hybrid scenarios, and finds that WhatIfBench remains far from saturated: even the strongest model reaches only a 64.62% final score.

Yu-Cheng Wang, Yuetian Du, Zheng Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Counterfactual Self-Evolving Agents for Evidence-Grounded Reasoning

Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verifi...

Xing Han, Yu-Xin Wang, Chen Chen et al. · 0 citations
Preprint Aug 2026

Counterfactual Explanations and the Scope of Contestability

The automation of consequential decisions through opaque machine learning models in societal domains impedes our agency. This paper is about how agency can be reinstated by the provision of certain kinds of knowledge. More precisely, we discuss whether a specific type of explanation, counterfactual explanations, facili...

Alice C. W. Huang, Thomas Grote · 0 citations
Conference Open access Sep 2026

What If. . . Counterfactual Explanations Were to Be Deployed?

Three lines of work are considered that build on standard counterfactual explanations and ask what would need to change for them to be better aligned with deployment requirements, including one that extends them beyond one-shot decisions to capture sequential decision-making.

Francesco Leofante · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.