Explanation audits reveal silent failures of machine learning models under distribution shift
TL;DR
These results support a scoped monitoring strategy for similar tabular settings: confidence-derived scores are effective for pointwise screening, whereas group-aware explanation audits provide complementary evidence about stable but incorrect feature reliance.
Abstract
Confidence-based monitoring is commonly used to decide when a machine-learning classifier should abstain, request human review, or raise an alert. However, a model may remain highly confident while relying on a spurious feature whose relationship with the target changes after deployment. We define such a silent failure as a substantial increase in realised error without a commensurate increase in a confidence-based warning signal. This study examines whether confidence-derived and explanation-derived monitoring signals identify complementary failure modes in controlled tabular settings. We evaluate logistic regression, random forests, LightGBM, and XGBoost on three real-data benchmarks subjected to additive noise, random masking, and mean shift, together with a matched synthetic shortcut-reversal benchmark. Pointwise monitoring compares confidence-complement uncertainty, predictive entropy, output volatility, explanation volatility, and a hybrid score, whereas batch-level auditing uses Jensen–Shannon attribution-profile drift and top-k feature drift. Across three repeated stratified splits, confidence-complement uncertainty achieved the strongest pointwise error ranking, with a mean AUROC of 0.8664, compared with 0.7524 for output volatility, 0.5752 for explanation volatility, and 0.7847 for the hybrid score. Predictive entropy produced an almost identical ranking (0.8664), as expected for binary classification. At the batch level, attribution-profile drift showed the highest mean Spearman correlation with realised error (0.6808), exceeding mean uncertainty (0.6132). Sensitivity analysis over \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$T\in \{3,5,10,20\}$$\end{document} perturbations confirmed that the weak pointwise performance of explanation volatility was not caused by the perturbation count. Under strong shortcut reversal, tree-based models made many high-confidence errors while both occlusion and TreeSHAP audits remained concentrated on the spurious feature group. These results support a scoped monitoring strategy for similar tabular settings: confidence-derived scores are effective for pointwise screening, whereas group-aware explanation audits provide complementary evidence about stable but incorrect feature reliance.