Skip to content
Open access

Machine learning approaches for diabetes prediction: A comparative analysis of classification models

2026 · Medical Science · 0 citations

TL;DR

These findings support the potential value of ensemble tree-based methods as a component of non-invasive diabetes risk-screening tools and highlight the continued relevance of glycaemic variables and anthropometric measures in T2DM risk stratification.

Abstract

Early and accurate identification of individuals at risk for type 2 diabetes mellitus (T2DM) is a clinical priority given the global scale of the epidemic. Machine learning (ML) methods offer a data-driven complement to conventional screening approaches; however, systematic comparisons of multiple classifiers under identical experimental conditions remain valuable for guiding model selection in practice. This study evaluated four ML classification algorithms — Naïve Bayes (NB), Logistic Regression (LR), Random Forest (RF), and Bagged Classification and Regression Trees (CART) — for T2DM prediction using the Pima Indians Diabetes Database (n = 768). Physiologically implausible zero values were treated as missing and median-imputed within each training fold, class weighting was applied to address the class imbalance between diabetic (34.9%) and non-diabetic (65.1%) participants, and hyperparameters were selected via grid search. Models were evaluated using 10-fold stratified cross-validation and reported as mean ± 95% confidence interval. Bagged CART achieved the highest overall accuracy (0.77 ± 0.04), sensitivity (0.78 ± 0.06), and F1-score (0.70 ± 0.05), while Naïve Bayes attained the highest specificity (0.82 ± 0.06). All models achieved comparable discrimination (AUC-ROC 0.81 – 0.84). SHAP-based feature importance analysis of the Bagged CART model identified plasma glucose concentration as the dominant predictor, followed by body mass index and age. These findings support the potential value of ensemble tree-based methods as a component of non-invasive diabetes risk-screening tools and highlight the continued relevance of glycaemic variables and anthropometric measures in T2DM risk stratification. As this analysis is retrospective and confined to a single, demographically restricted cohort, these results should be regarded as hypothesis-generating rather than evidence of clinical readiness, pending external validation.

Read PDF

Similar papers

Open access 2026

A Comparative Evaluation of Various Machine Learning Techniques for Prediction of Type 2 Diabetes Mellitus

This research uses Machine Learning algorithms to evaluate the likelihood of Type 2 Diabetes Mellitus (T2DM) using lifestyle and family history data and presents a performance assessment of seven ML classifiers, showing that SVM achieved the highest accuracy and AUC, while LR, RF, and XGBoost also performed competitive...

Rizwan Akhtar, Muhammad Kalamuddin Ahamad · 0 citations
Open access Aug 2026

A Comparative Evaluation of Machine Learning Algorithms for Diabetes Risk Prediction

Evaluated machine learning algorithms for predicting diabetes risk from routinely available clinical and lifestyle variables confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained setti...

T. Olayinka · 0 citations

Diabetes Prediction Using Machine Learning Model: A comparative Approach

Six supervised learning models were developed and compared for diabetes prediction using a dataset and compared for diabetes prediction using a 100k patients records with eight clinical features including gender, age, hypertension, smoking history, heart disease, BMI, HbA1c level, and blood glucose level.

Akshay Bhardwaj, Rajesh Chauhan, Devansh Khajuria · 0 citations
Open access 2026

A Comparative Analysis of Machine Learning Algorithms for the Early Prediction of Diabetes with an Evaluation of Class-Imbalance Handling

The study comes to the conclusion that headline accuracy is an unreliable guide in imbalanced medical prediction, that imbalance handling can change a model's practical usefulness, and that this benefit is strongly algorithm-dependent, meaning that the decision to resample should be based on the algorithm and the scree...

A. Oduroye, Temilade Opanuga, Esther Tosin Akanbi et al. · 0 citations
Open access Sep 2026

MACHINE LEARNING APPROACHES FOR EARLY DIABETES PREDICTION: A COMPARATIVE STUDY USING CLINICAL DATA

This study focuses on the accurate and early prediction of diabetes, which plays a vital role in improving clinical care and treatment planning. The research evaluates the performance of three machine learning techniques - Logistic Regression, Random Forest, and Support Vector Machine (SVM) - using data collected from...

N. Sri Ram, P. Arumugam, M. G. · 0 citations
Review Open access Aug 2026

Machine Learning for Diabetes Screening: Insights from the Behavioral Risk Factor Surveillance System and Pima Indians Diabetes Databases

A simple logistic regression – with excellent recall and competitive AUC – appears preferable for population screening, where transparency, robustness to reporting noise, and ease of implementation are paramount.

Shaked Shechter, N. Shomron · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.