Explainable Machine Learning for Type 2 Diabetes Screening Using Shap Feature Attribution on NHANES 2017-2018 Data
Abstract
Type 2 Diabetes Mellitus (T2DM) presents a critical public health challenge, particularly in Southeast Asian lowand middle-income countries where healthcare resources are constrained. This paper evaluates and compares two machine learning classifiers - Random Forest (RF) and XGBoost - for T2DM risk classification using the NHANES 2017-2018 dataset (5.393 adult participants). SHAP (SHapley Additive exPlanations) is applied to both models to provide clinically interpretable feature attribution. XGBoost achieved the highest overall performance with accuracy of 91.84%, precision of 0.8372, F1-score of 0.7105, and AUC-ROC of 0.929. SHAP analysis consistently identified HbA1c, age, and waist circumference as dominant predictors across both models. This work constitutes the ML classification and explainability phase of a broader programme toward an Explainable AI-Driven Digital Twin Framework for T2DM management in Southeast Asian health information systems; Digital Twin architecture and HL7 FHIR integration are reserved for subsequent phases.