Aug 2026· Academic Journal of International University of Erbil· Vol 3, pp. 442-451· 0 citations· 15 references
TL;DR
The results prove that ensemble models applied to class balancing and explainability methods can be used to rely on and provide clear tools regarding multiclass diabetes classification and illustrate the costs and benefits of operating a global model versus a minority-class sensitive one.
Abstract
Proper identification of the stages of diabetes (non-diabetic, prediabetic, diabetic) is pivotal in the early intervention and risk stratification. Nevertheless, class imbalance in clinical datasets often imbalances machine learning models in the context of majority classes. In this paper, we have assessed four classifiers namely, the Support Vector Machine (SVM), Artificial Neural Network (ANN) and the Logistic Regression (LR) and on a multiclass dataset on diabetes. Once the duplicate cases of patients were eliminated 264 distinct cases were retained. SMOTE was only used on the training data to overcome the problem of class imbalance. Findings indicate that Random Forest performed the highest with an accuracy of 98.1, F1-score of 98.5, and AUC of 99.9 and stability before and after balancing. Logistic Regression and ANN demonstrated some significant improvements in recall following SMOTE, which illustrates the costs and benefits of operating a global model versus a minority-class sensitive one. SHAP analysis found that HbA1c, age, LDL, and BMI are the most influential predictors that improve interpretability and clinical relevance. These results prove that ensemble models applied to class balancing and explainability methods can be used to rely on and provide clear tools regarding multiclass diabetes classification.
The study comes to the conclusion that headline accuracy is an unreliable guide in imbalanced medical prediction, that imbalance handling can change a model's practical usefulness, and that this benefit is strongly algorithm-dependent, meaning that the decision to resample should be based on the algorithm and the scree...
A. Oduroye, Temilade Opanuga, Esther Tosin Akanbi et al.· International journal of re...· 0 citations
The pipeline approach ensured no data leakage in cross-validation, and the findings support ensemble ML models with SMOTE as a preprocessing step for imbalanced CVD datasets.
M. Maindarkar· Journal of Intelligent Decis...· 0 citations
Diabetes mellitus is a widespread metabolic disorder marked by chronic hyperglycemia and severe complications. Early and accurate detection is crucial for effective management and preventing disease progression. This study systematically evaluates the performance of three ensemble learning strategies Bagging, Boostin...
Evaluated machine learning algorithms for predicting diabetes risk from routinely available clinical and lifestyle variables confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained setti...
T. Olayinka· FUDMA Journal of Sciences· 0 citations
Comparisons of the performance of the Random Forest and Support Vector Machine algorithms in predicting diabetes and the effect of applying the Synthetic Minority Over-sampling Technique to imbalanced data show that Random Forest outperforms SVM.
Baharudin Yusuf· Jurnal Informatika dan Tekni...· 0 citations
Six supervised learning models were developed and compared for diabetes prediction using a dataset and compared for diabetes prediction using a 100k patients records with eight clinical features including gender, age, hypertension, smoking history, heart disease, BMI, HbA1c level, and blood glucose level.