These findings support the potential value of ensemble tree-based methods as a component of non-invasive diabetes risk-screening tools and highlight the continued relevance of glycaemic variables and anthropometric measures in T2DM risk stratification.
Abstract
Early and accurate identification of individuals at risk for type 2 diabetes mellitus (T2DM) is a clinical priority given the global scale of the epidemic. Machine learning (ML) methods offer a data-driven complement to conventional screening approaches; however, systematic comparisons of multiple classifiers under identical experimental conditions remain valuable for guiding model selection in practice. This study evaluated four ML classification algorithms — Naïve Bayes (NB), Logistic Regression (LR), Random Forest (RF), and Bagged Classification and Regression Trees (CART) — for T2DM prediction using the Pima Indians Diabetes Database (n = 768). Physiologically implausible zero values were treated as missing and median-imputed within each training fold, class weighting was applied to address the class imbalance between diabetic (34.9%) and non-diabetic (65.1%) participants, and hyperparameters were selected via grid search. Models were evaluated using 10-fold stratified cross-validation and reported as mean ± 95% confidence interval. Bagged CART achieved the highest overall accuracy (0.77 ± 0.04), sensitivity (0.78 ± 0.06), and F1-score (0.70 ± 0.05), while Naïve Bayes attained the highest specificity (0.82 ± 0.06). All models achieved comparable discrimination (AUC-ROC 0.81 – 0.84). SHAP-based feature importance analysis of the Bagged CART model identified plasma glucose concentration as the dominant predictor, followed by body mass index and age. These findings support the potential value of ensemble tree-based methods as a component of non-invasive diabetes risk-screening tools and highlight the continued relevance of glycaemic variables and anthropometric measures in T2DM risk stratification. As this analysis is retrospective and confined to a single, demographically restricted cohort, these results should be regarded as hypothesis-generating rather than evidence of clinical readiness, pending external validation.
This research uses Machine Learning algorithms to evaluate the likelihood of Type 2 Diabetes Mellitus (T2DM) using lifestyle and family history data and presents a performance assessment of seven ML classifiers, showing that SVM achieved the highest accuracy and AUC, while LR, RF, and XGBoost also performed competitive...
Rizwan Akhtar, Muhammad Kalamuddin Ahamad· ITEGAM- Journal of Engineeri...· 0 citations
Evaluated machine learning algorithms for predicting diabetes risk from routinely available clinical and lifestyle variables confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained setti...
T. Olayinka· FUDMA Journal of Sciences· 0 citations
Six supervised learning models were developed and compared for diabetes prediction using a dataset and compared for diabetes prediction using a 100k patients records with eight clinical features including gender, age, hypertension, smoking history, heart disease, BMI, HbA1c level, and blood glucose level.
The study comes to the conclusion that headline accuracy is an unreliable guide in imbalanced medical prediction, that imbalance handling can change a model's practical usefulness, and that this benefit is strongly algorithm-dependent, meaning that the decision to resample should be based on the algorithm and the scree...
A. Oduroye, Temilade Opanuga, Esther Tosin Akanbi et al.· International journal of re...· 0 citations
This study focuses on the accurate and early prediction of diabetes, which plays a vital role in improving clinical care and treatment planning. The research evaluates the performance of three machine learning techniques - Logistic Regression, Random Forest, and Support Vector Machine (SVM) - using data collected from...
N. Sri Ram, P. Arumugam, M. G.· International Journal of Eng...· 0 citations
A simple logistic regression – with excellent recall and competitive AUC – appears preferable for population screening, where transparency, robustness to reporting noise, and ease of implementation are paramount.
Shaked Shechter, N. Shomron· Journal of Scientific Innova...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.