Skip to content
Open access

Interpretable machine learning for predicting mild cognitive impairment in elderly patients with type 2 diabetes mellitus: model development and performance assessment.

Jul 2026 · Scientific Reports · 0 citations
Medicine

TL;DR

An interpretable RF-based model using routine clinical data effectively predicted MCI in elderly patients with T2DM and provided clinically intuitive explanations of risk drivers, and may support risk-stratified cognitive screening and individualized management.

Abstract

Type 2 diabetes mellitus (T2DM) is a prevalent chronic condition, particularly in the elderly, and is associated with an increased risk of cognitive decline, including mild cognitive impairment (MCI). This study aimed to develop and validate an interpretable machine learning (IML) model to predict MCI in elderly T2DM patients using routine clinical data. A retrospective cohort of 923 elderly T2DM patients (≥ 60 years) was analyzed, with data collected from January 2021 to January 2025. Key MCI predictors were selected using a two-stage feature selection process involving Boruta and least absolute shrinkage and selection operator (LASSO). Six machine learning (ML) algorithms-logistic regression (LR), extreme gradient boosting (XGBoost), support vector machine (SVM), k-nearest neighbor (KNN), random forest (RF), and decision tree (DT)-were trained and evaluated. SHapley Additive exPlanations (SHAP) were employed to interpret model predictions and provide insights into feature importance. Among 923 participants, 424 (45.9%) had MCI. Baseline characteristics were comparable between the training and validation sets, with similar MCI prevalence (46.1% vs. 45.5%). Seven predictors were consistently selected: age, years of education, duration of diabetes, regular physical activity, cerebrovascular disease, glycated hemoglobin (HbA1c), and fasting plasma glucose (FPG). Among the six models, the RF model demonstrated the best overall performance, achieving an AUC of 0.842 in the validation set, with favorable discrimination, calibration, and net clinical benefit. SHAP analysis identified duration of diabetes as the most influential predictor, followed by HbA1c and FPG, emphasizing the role of both cumulative and current glycemic burden. An interpretable RF-based model using routine clinical data effectively predicted MCI in elderly patients with T2DM and provided clinically intuitive explanations of risk drivers. This approach may support risk-stratified cognitive screening and individualized management; external multicenter validation and prospective evaluation are warranted.

Read PDF

Similar papers

Open access Jul 2026

An interpretable machine learning model for predicting 1-year major adverse cardiovascular events in patients with type 2 diabetes and hypertension

Background Patients with coexisting type 2 diabetes mellitus (T2DM) and hypertension (HTN) face a synergistically elevated risk of major adverse cardiovascular events (MACE). Evidence for prediction models developed specifically in established T2DM-HTN comorbidity population remains limited. Objective To methodologically explore and preliminarily evaluate an interpretable machine learning framework for 1-year MACE prediction in hospitalized patients with coexisting T2DM and HTN using routine clinical data. Methods This retrospective study included 1,054 hospitalized patients with T2DM and HTN, of whom 249 (23.6%) experienced MACE during 1-year follow-up. The dataset was randomly divided into training (60%), validation (20%), and independent test (20%) cohorts using stratified sampling. LASSO regression was applied for feature selection from 69 clinical variables. Four algorithms, including logistic regression, random forest, support vector machine, and XGBoost, were developed and compared. Model performance was assessed using discrimination, calibration, and clinical utility metrics. SHapley Additive exPlanations (SHAP) were used to interpret the final model. Results LASSO identified six stable predictors: HbA1c, age, hypertension duration, cystatin C (CysC), T2DM duration, and carotid intima-media thickness (CIMT). Sex was additionally incorporated based on clinical relevance. Multivariable logistic regression showed that HbA1c, age, hypertension duration, T2DM duration, CysC, and CIMT were associated with 1-year MACE risk, whereas sex was not statistically significant. Logistic regression showed the best relative balance between discrimination, calibration, and simplicity on the validation set, although learning curves indicated limited incremental improvement with increasing training sample size. After isotonic regression recalibration, the final logistic regression model achieved an ROC-AUC of 0.828, a PR-AUC of 0.656, and a Brier score of 0.116 on the independent test set. Decision curve analysis indicated potential clinical net benefit. SHAP linked model predictions to glycemic burden, aging, cumulative disease exposure, renal-related risk, and subclinical atherosclerosis. Conclusion An interpretable logistic regression model based on seven routine clinical variables showed relatively good internal performance for predicting 1-year composite MACE risk in hospitalized patients with coexisting T2DM and HTN. CysC provided additional prognostic information beyond its conventional role as a renal filtration marker, although this association should be interpreted as prognostic rather than causal. External validation is required before the model can be considered for clinical decision support.

Juan Lv, Xi-Rui Wang, Zhengyi Zhang · 0 citations
Open access 2026

A Comparative Evaluation of Various Machine Learning Techniques for Prediction of Type 2 Diabetes Mellitus

Diabetes affects over 101 million people in India, with many more at risk due to routine and hereditary factors. Early diagnosis is crucial to prevent complications, which make accurate predictive tools essential in healthcare. This research uses Machine Learning (ML) algorithms to evaluate the likelihood of Type 2 Diabetes Mellitus (T2DM) using lifestyle and family history data. The trained models demonstrate strong predictive ability, allowing individuals to self-assess their risk and supporting healthcare professionals in early detection and intervention. This study presents a performance assessment of seven ML classifiers: Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), Logistic Regression (LR), Naïve Bayes (NB), k-Nearest Neighbor (k-NN), and Extreme Gradient Boosting (XGBoost). These classifiers were applied to the widely used PIMA Indian Diabetes dataset (PIDD), which contains 768 clinical records of adult women aged 21 and above, providing key medical information for diabetes analysis. Multiple evaluation measures were applied to assess model performance with results showing that SVM achieved the highest accuracy and AUC, while LR, RF, and XGBoost also performed competitively. Although k-NN attained the highest recall, it yielded a higher false positive rate. These findings highlight that no single model is perfect for every situation, and the choice of classifier should match clinical needs. This study serves as a reference for ML applications in diabetes prediction.

Rizwan Akhtar, Muhammad Kalamuddin Ahamad · 0 citations
Open access Jul 2026

Development and validation of an interpretable machine learning model for predicting atrial fibrillation risk in middle-aged and older patients with coronary heart disease

This data-driven, interpretable XGBoost model enables individualized AF risk assessment in middle-aged and older CHD patients, offering a practical tool for early identification and targeted intervention in clinical practice.

Feng Chen, Qin Fu, Ling Li et al. · 0 citations
Open access Aug 2026

An interpretable machine learning screening model for MoCA-defined possible mild cognitive impairment in rural Xinjiang: a preliminary framework for resource-limited primary care

To construct and validate an interpretable machine learning screening model for MoCA-defined possible mild cognitive impairment applicable to residents in rural areas of Xinjiang, China. A total of 708 residents from rural Xinjiang were recruited between June and July 2025. Six machine learning methods—Logistic Regression (LR), Adaptive Boosting (AdaBoost), Multilayer Perceptron (MLP), Naïve Bayes (NB), Extreme Gradient Boosting (XGBoost), and Light Gradient Boosting Machine (LightGBM)—were employed to identify MoCA-defined possible MCI based on data from four dimensions: physiological, psychological, social, and behavioral. The optimal model was selected using the Area Under the Curve (AUC) as the primary evaluation metric. Model interpretability was assessed using SHapley Additive explanations (SHAP), and the dose–response relationships between continuous variables and MoCA-defined possible MCI were visualized using Restricted Cubic Splines (RCS). After feature selection, nine key variables were retained for model construction. Among the six models developed, the XGBoost model demonstrated the best performance, achieving an AUC of 0.839 (95% CI: 0.811–0.867) in the training set and 0.747 (95% CI: 0.673–0.820) in the validation set. SHAP analysis revealed that Education, Age, and Direct Bilirubin (DBIL) were the three most influential predictors. Restricted cubic spline analysis indicated linear correlations with MoCA-defined possible MCI for Education, Age, DBIL, and Systolic Blood Pressure (SBP) (overall p  < 0.05, non-linear p  > 0.05). A non-linear association was observed for Triglycerides (TG) (overall p  < 0.001, non-linear p  < 0.001), with its dose–response curve exhibiting a complex non-linear trend. We developed and validated an interpretable screening model for MoCA-defined possible MCI tailored to rural Xinjiang by integrating established machine learning techniques within a localized framework. The primary contribution of this work lies in the applied and translational validation of these methods for an understudied, resource-limited population rather than in algorithmic innovation. The finalized XGBoost model requires only nine easily obtainable variables, offering moderate discriminative performance and useful interpretability for preliminary screening purposes. This preliminary screening framework may offer a potentially useful approach for cognitive risk stratification in resource-limited primary care settings, though independent external validation is required before broader implementation.

Yu-Tong Li, Lili He, Jia-Huan He et al. · 0 citations
Open access Aug 2026

Development and Validation of a Disability Risk Prediction Model for Older Adults Based on Machine Learning: A Multi-Algorithm Comparison with SHAP Interpretation

Objective To develop a predictive model for disability risk in older adults using machine learning algorithms. Methods A convenience sample of 13,809 older adults (aged ≥60 years) was recruited from seven medical institutions, three communities, and five nursing homes in Zunyi City, Guizhou Province. Participants were randomly divided into a training set (n = 9667) and a validation set (n = 4142) at a 7:3 ratio. Disability status was used as the outcome variable. Nine machine learning algorithms—logistic regression, decision tree, random forest, XGBoost, LightGBM, support vector machine, artificial neural network, K‑nearest neighbor, and naïve Bayes—were used to construct prediction models. Model performance was evaluated using area under the receiver operating characteristic curve (AUC), accuracy, and other metrics, and the best‑performing model was selected. The SHapley Additive exPlanations (SHAP) method was used for interpretability analysis of the optimal model. Results Among the 13,809 participants, 5308 (38.44%) were identified as having disability. Among the nine models, LightGBM achieved the highest AUC (0.859), accuracy (0.792), precision (0.771), sensitivity (0.651), specificity (0.880), and F1 score (0.706). Conclusion Among the developed prediction models, the LightGBM‑based model demonstrated superior overall predictive performance in internal validation, providing a reference for disability management in older adults.

Shaoting Yang, Heting Liang, Yamin Peng et al. · 0 citations
Open access Aug 2026

Machine learning-based prediction of prolonged length of stay in older patients with type 2 diabetes mellitus and cardiovascular disease

Background Older patients with type 2 diabetes mellitus (T2DM) and cardiovascular disease (CVD) frequently experience prolonged length of stay (PLOS). This condition increases healthcare burden and worsens prognosis. However, no predictive model specifically addresses PLOS in this high-risk multimorbid population. Methods This single-center retrospective study included hospitalized older T2DM-CVD patients. PLOS was defined as hospital stay exceeding the 75th percentile of the training set population. Potential predictors were selected via LASSO regression. Eight machine learning (ML) models were developed to predict PLOS risk. Model performance was evaluated using receiver operating characteristic curves, calibration curves, and decision curve analysis. SHAP analysis was employed for model interpretability. Results A total of 27,629 patients were included. The XGBoost model achieved the highest training AUC (0.819) and demonstrated competitive predictive performance in both the internal (AUC = 0.753) and time-based external (AUC = 0.728) validation sets. However, its performance advantage over simpler models such as logistic regression was modest in validation, and XGBoost showed some degree of overfitting (AUC drop of 0.066 from training to validation). Although logistic regression showed comparable validation performance with less overfitting, XGBoost was selected as the final model for its ability to capture complex nonlinear interactions and provide SHAP-based interpretability, with the understanding that further external validation is needed. Key predictors included cerebral infarction, white blood cell count, anemia, pulse rate, the glycated hemoglobin to high-density lipoprotein cholesterol ratio (GHR), and osteoporosis. Most continuous variables showed nonlinear associations with PLOS risk. Conclusions The XGBoost-based model effectively predicts PLOS risk in older T2DM-CVD patients. This tool shows promise for early identification of high-risk individuals and optimization of medical resource allocation within our institutional setting. However, further prospective and multi-center validation studies are required before clinical adoption.

Yixia Zuo, Jianfei Chen, Jie Wang et al. · 0 citations