An interpretable machine learning screening model for MoCA-defined possible mild cognitive impairment in rural Xinjiang: a preliminary framework for resource-limited primary care
Abstract
To construct and validate an interpretable machine learning screening model for MoCA-defined possible mild cognitive impairment applicable to residents in rural areas of Xinjiang, China. A total of 708 residents from rural Xinjiang were recruited between June and July 2025. Six machine learning methods—Logistic Regression (LR), Adaptive Boosting (AdaBoost), Multilayer Perceptron (MLP), Naïve Bayes (NB), Extreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM)—were employed to identify MoCA-defined possible MCI based on data from four dimensions: physiological, psychological, social, and behavioral. The optimal model was selected using the Area Under the Curve (AUC) as the primary evaluation metric. Model interpretability was assessed using SHapley Additive explanations (SHAP), and the dose–response relationships between continuous variables and MoCA-defined possible MCI were visualized using Restricted Cubic Splines (RCS). After feature selection, nine key variables were retained for model construction. Among the six models developed, the XGBoost model demonstrated the best performance, achieving an AUC of 0.839 (95% CI: 0.811–0.867) in the training set and 0.747 (95% CI: 0.673–0.820) in the validation set. SHAP analysis revealed that Education, Age, and Direct Bilirubin (DBIL) were the three most influential predictors. Restricted cubic spline analysis indicated linear correlations with MoCA-defined possible MCI for Education, Age, DBIL, and Systolic Blood Pressure (SBP) (overall p < 0.05, non-linear p > 0.05). A non-linear association was observed for Triglycerides (TG) (overall p < 0.001, non-linear p < 0.001), with its dose–response curve exhibiting a complex non-linear trend. We developed and validated an interpretable screening model for MoCA-defined possible MCI tailored to rural Xinjiang by integrating established machine learning techniques within a localized framework. The primary contribution of this work lies in the applied and translational validation of these methods for an understudied, resource-limited population rather than in algorithmic innovation. The finalized XGBoost model requires only nine easily obtainable variables, offering moderate discriminative performance and useful interpretability for preliminary screening purposes. This preliminary screening framework may offer a potentially useful approach for cognitive risk stratification in resource-limited primary care settings, though independent external validation is required before broader implementation.