Auditing Population-Level XAI Agreement with cABC: Evidence from Diabetes Risk Prediction
Abstract
Auditing agreement between global explanation methods is underdeveloped in clinical XAI. Standard population-level agreement measures are poorly aligned with the practical question of interest: top-K overlap depends on an arbitrary cutoff, while rank-correlation metrics can overweight tail-order differences that are operationally negligible. We address this problem by combining computed ABC (cABC) analysis with per-group Jaccard indices to compare feature-importance rankings at the level of data-driven importance groups rather than raw rank positions. We apply this framework to SHAP and Permutation Importance (PI) in diabetes risk prediction using three Behavioral Risk Factor Surveillance System (BRFSS) cohorts (2015: n = 253,680; 2021: n = 236,378; 2023: n = 272,769) and three classifiers: XGBoost, Random Forest, and Logistic Regression. Across nine model–cohort settings, cross-method agreement between SHAP and PI is high, with mean Group-A Jaccard JA =0.9127, whereas cross-model agreement is lower, with mean JA =0.8307. This suggests that, in our experimental setting, the choice of model pipeline may perturb global feature-importance structure more than attribution-method choice. Five features (Age, BMI, GenHlth, HighBP, and HighChol) remain essential across all 18 attribution rankings. These results position cABC-based group-overlap analysis as a data-driven framework for population-level XAI auditing. In this setting, the framework reveals a practically relevant observation for diabetes risk prediction: stable identification of essential predictors is more sensitive to model-pipeline choice than to whether explanations are generated with SHAP or PI. The source code and dataset are available at https://github.com/thieuanhvan/diabetes-xai-agreement.