Skip to content
Conference Open access

Beyond Top-1: Addressing Inconsistencies in Evaluating Counterfactual Explanations for Recommender Systems (Extended Abstract)

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · 6 citations · 38 references

Abstract

Counterfactual explanations have become an important paradigm for improving the transparency of machine learning models by showing how small input changes can alter model outputs. While substantial progress has been made in generating such explanations, their evaluation remains insufficiently standardized, particularly for systems that produce ranked outputs rather than single-label predictions. Existing evaluation protocols in recommender systems commonly focus only on whether the top-1 recommendation changes after perturbation. We argue that this practice can lead to inconsistent and misleading conclusions, as the relative ranking of explanation methods may vary with changes in the quality of the underlying model. In this work, we revisit the evaluation of counterfactual explanations from a ranking perspective. We propose extending top-1 evaluation to list-wise top-k protocols that assess explanation effectiveness across multiple highly ranked outputs. Through experiments on multiple datasets, recommender architectures, and explanation methods, we show that top-k evaluation substantially improves consistency and yields more reliable comparisons between competing explainers. Our findings highlight a broader methodological lesson for explainable AI: when models return ranked results, explanation quality should be assessed using ranking-aware evaluation protocols rather than top-1 criteria alone.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.