This work introduces CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits, enabling us to evaluate and improve methods for explaining LLM behaviors.
Abstract
Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a"good"explanation? In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs. To this end, we introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits. This yields thousands of high-quality explanations for naturally-occurring model behaviors along with supporting counterfactual evidence. We apply CHIVE in two ways. First, we evaluate whether common LLM interpretability techniques improve an agent's ability to predict counterfactual model behaviors. Surprisingly, we find no uplift from any of the interpretability techniques studied. Second, we use CHIVE to generate training data. We find that training models to predict outcomes of CHIVE-generated counterfactual experiments generalizes to various out-of-distribution settings. Overall, CHIVE automatically discovers explanations of naturally-occurring LLM behaviors, enabling us to evaluate and improve methods for explaining LLM behaviors.
This work evaluates ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales and shows that model scale is the strongest determinant of explanation quality.
Giannis Kalyvas, Giorgos Filandrianos, Orfeas Menis-Mastromichalakis et al.· 0 citations
It is found that an organization's choices on measurement models for feature and labels, business requirements, model validation, and the metric of model success have as much or more impact on the generated counterfactuals as the specifics of the generating method.
Mattia Cerrato, Otto Sahlgren, Xenia Heilmann· 0 citations
A diagnostic benchmark for open-domain, open-form, long-horizon counterfactual causal reasoning, containing 220 what-if questions across STEM, HSS, and Hybrid scenarios, and finds that WhatIfBench remains far from saturated: even the strongest model reaches only a 64.62% final score.
Yu-Cheng Wang, Yuetian Du, Zheng Liu et al.· 0 citations
Self-play proposer--solver methods improve reasoning by generating tasks and learning from verified solutions. However, for evidence-identifiable tasks, where case-specific evidence and domain knowledge determine a checkable answer, self-play requires generating plausible cases whose answers can be independently verifi...
Xing Han, Yu-Xin Wang, Chen Chen et al.· 0 citations
The automation of consequential decisions through opaque machine learning models in societal domains impedes our agency. This paper is about how agency can be reinstated by the provision of certain kinds of knowledge. More precisely, we discuss whether a specific type of explanation, counterfactual explanations, facili...
Three lines of work are considered that build on standard counterfactual explanations and ask what would need to change for them to be better aligned with deployment requirements, including one that extends them beyond one-shot decisions to capture sequential decision-making.
Francesco Leofante· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.