Skip to content

CEL: Comprehensive Counterfactual Explanations Library and Benchmark

Jul 2026 · arXiv.org · Vol abs/2607.22045 · 0 citations · 33 references
Computer Science

TL;DR

CEL (Counterfactual Explanations Library), a unified library and benchmark for counterfactual explanations designed to support consistent implementation and evaluation, is introduced, believed to be the first comprehensive benchmark that systematically evaluates recent counterfactual explanation methods within a unified and reproducible framework.

Abstract

Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome. While early methods primarily focused on minimal feature changes, recent work incorporates additional properties such as sparsity, actionability and plausibility. Despite this progress, fair and systematic evaluation remains challenging. Existing studies often rely on different data splits, predictive models, and evaluation metrics, which limits objective comparison across methods. To fill this gap, we introduce CEL (Counterfactual Explanations Library), a unified library and benchmark for counterfactual explanations designed to support consistent implementation and evaluation. CEL includes 18 datasets of varying size and complexity and provides implementations or reimplementations of 14 widely used counterfactual methods. Using this standardized setup, we conduct a comprehensive quantitative comparison across a variety of methods on datasets that differ in size, number, and types of attributes. The evaluation protocol incorporates multiple complementary metrics capturing validity, coverage, sparsity, proximity, and distributional plausibility, including density- and outlier-based measures to assess the realism of generated counterfactuals. To the best of our knowledge, this is the first comprehensive benchmark that systematically evaluates recent counterfactual explanation methods within a unified and reproducible framework. While prior libraries and benchmarking efforts exist in the literature, many are outdated, limited in scope, or lack consistent evaluation protocols. The proposed benchmark aims to improve reproducibility, enable fair comparison, and establish a workbench for the development of future counterfactual explanation methods.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Robust to Which Model Change? A Unified Evaluation of Robust Counterfactual Explanations

Robust counterfactual explanations promise recourse that still works after the model behind it changes. Whether they keep that promise depends on what the change is. A small perturbation of the parameters, retraining on new data, and a new architecture are different events, and each existing method is evaluated against...

Marcin Kostrzewa, Maciej Ziȩba · 0 citations
#artificial intelligence Preprint Aug 2026

Taking the Whys Seriously: Limitations of Counterfactual Explanations in Justification and Recourse

It is found that an organization's choices on measurement models for feature and labels, business requirements, model validation, and the metric of model success have as much or more impact on the generated counterfactuals as the specifics of the generating method.

Mattia Cerrato, Otto Sahlgren, Xenia Heilmann · 0 citations
#artificial intelligence Preprint Sep 2026

AVCG: A Generalized Variational Framework for Counterfactual Generation under Hypothesis Distributions

Counterfactual explanations formalize"what-if"scenarios by identifying modifications to an input instance that obtain a desired alternative prediction. Traditionally, whether generated via instance-specific optimization or amortized single pass models, these approaches rely on a single, deterministic point-estimate pre...

J. Duell, Alejandro Jimenez Rodriguez, Mahault Albarracin · 0 citations
#machine learning Preprint Sep 2026

FCx: An algorithm for finding Feasible Counterfactual Explanations

Counterfactual (CF) explanations identify changes that alter an input's classification. While existing methods produce realistic and low-cost CFs, they often fail to ensure feasibility, by suggesting non-constructive modifications or incompatible with future changes (e.g., changing an individual's race to secure a job...

Kleopatra Markou, V. Kalogeraki, Dimitrios Gunopulos · 0 citations
Open access Aug 2026

Measuring What Matters: Consistency and Compactness in Evaluation of Counterfactual Explanations

Explainability in recommender systems (RS) remains a pivotal challenge. Counterfactual explanations have emerged as a particularly actionable paradigm, offering intuitive “what-if” reasoning. However, their evaluation lacks principled standards. Current metrics primarily assess whether explanations change the top-ranke...

Amir Reza Mohammadi, Andreas Peintner, Michael M. Müller et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.