Skip to content
Review Open access

Explanation audits reveal silent failures of machine learning models under distribution shift

Aug 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 39 references
Computer Science

TL;DR

These results support a scoped monitoring strategy for similar tabular settings: confidence-derived scores are effective for pointwise screening, whereas group-aware explanation audits provide complementary evidence about stable but incorrect feature reliance.

Abstract

Confidence-based monitoring is commonly used to decide when a machine-learning classifier should abstain, request human review, or raise an alert. However, a model may remain highly confident while relying on a spurious feature whose relationship with the target changes after deployment. We define such a silent failure as a substantial increase in realised error without a commensurate increase in a confidence-based warning signal. This study examines whether confidence-derived and explanation-derived monitoring signals identify complementary failure modes in controlled tabular settings. We evaluate logistic regression, random forests, LightGBM, and XGBoost on three real-data benchmarks subjected to additive noise, random masking, and mean shift, together with a matched synthetic shortcut-reversal benchmark. Pointwise monitoring compares confidence-complement uncertainty, predictive entropy, output volatility, explanation volatility, and a hybrid score, whereas batch-level auditing uses Jensen–Shannon attribution-profile drift and top-k feature drift. Across three repeated stratified splits, confidence-complement uncertainty achieved the strongest pointwise error ranking, with a mean AUROC of 0.8664, compared with 0.7524 for output volatility, 0.5752 for explanation volatility, and 0.7847 for the hybrid score. Predictive entropy produced an almost identical ranking (0.8664), as expected for binary classification. At the batch level, attribution-profile drift showed the highest mean Spearman correlation with realised error (0.6808), exceeding mean uncertainty (0.6132). Sensitivity analysis over \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$T\in \{3,5,10,20\}$$\end{document} perturbations confirmed that the weak pointwise performance of explanation volatility was not caused by the perturbation count. Under strong shortcut reversal, tree-based models made many high-confidence errors while both occlusion and TreeSHAP audits remained concentrated on the spurious feature group. These results support a scoped monitoring strategy for similar tabular settings: confidence-derived scores are effective for pointwise screening, whereas group-aware explanation audits provide complementary evidence about stable but incorrect feature reliance.

Read PDF

Similar papers

Open access 2026

Explainable Anomaly Detection in Accounting Journal Entries Using Autoencoders and Counterfactual Explanations

In financial auditing, an autoencoder (AE) can flag a journal entry as anomalous, but the auditor still needs to know which attributes triggered the flag and how to correct the entry. We address both questions with a pipeline that uses attribute-level SHAP (SHapley Additive exPlanations) attribution (RESHAPE) for root-...

Y. Kawazura, Daisuke Saitou · 0 citations
Open access Aug 2026

Reliability auditing of explanations for machine-learning-based intrusion detection systems

Machine-learning-based intrusion detection systems can learn nonlinear and interaction-based traffic patterns that are difficult to capture using static rules, but their predictions remain difficult to interpret reliably in analyst-facing cybersecurity workflows. This paper proposes a unified quantitative framework for...

Elijah M. Maseno, Yan-Xia Sun, Zeng-Hui Wang · 0 citations
Review Open access Sep 2026

Explainable and Analyst-Driven Random Forest for Intrusion Detection

Random Forest and other tree-ensemble classifiers achieve high accuracy in network intrusion detection; however, their aggregate decision logic prevents analysts from auditing or deploying individual predictions as operational rules. Post hoc explanation methods introduce latencies incompatible with security operation...

Saloua Bellouch, Mostapha Zbakh, S. Aouad et al. · 0 citations

Expert Systems With Applications

LLR2, a novel fully unsupervised concept drift detector based on Bayesian networks, is proposed, which computes a log-likelihood ratio that is evaluated against chi-squared quantiles for each variable.

Rafael Sojo, C. Bielza, P. Larrañaga · 0 citations
#artificial intelligence Review Aug 2026

RiskBlend: A Multi-Signal Framework for Test Input Prioritization in Machine Learning Regression Testing

RiskBlend is proposed, a classifier-agnostic prioritization framework that combines four complementary risk signals: historical failure patterns, prediction shift, decision-boundary shift, and neighborhood change that achieves the highest average APFD in all 80 dataset-classifier-scenario combinations.

Madhusudan Srinivasan, Namith Nishal Raphae · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.