Skip to content
Open access

Optimizing diabetes prediction in machine learning models: Evaluating the effectiveness of a novel class imbalance technique—adaptive synthetic class balancing with class proportion filtering

2026 · Sigma Journal of Engineering and Natural Sciences · 0 citations · 35 references

TL;DR

Adaptive Synthetic Class Balancing with filtering by class ratio (ASCPF) is suggested, a combination of creating smart fake samples and ratio control which results in more uniform performance on classes with skewed distribution, especially in important health decisions.

Abstract

The skewedness of results when predicting diabetes is mostly due to uneven distribution of data, especially in reducing detection rates of real patients. These are errors which cause delay in treatment or incorrect diagnosis. This work suggests a counter plan to this assumption, which is Adaptive Synthetic Class Balancing with filtering by class ratio (ASCPF). This does not require siloing such methods as SMOTE or NearMiss, but it does change the way to create or make samples, as well as the way to achieve retention, depending on group size. We evaluate the performance of each strategy by considering rare cases, using Extra Trees Classifier. We ran tests on the following data collections: CDC Diabetes, Breast Cancer Wisconsin (Diagnostic), and Credit Card Fraud Detection, KDD Cup 1999 Intrusion Detection, PIMA, and BRFSS. In the performance trend, ASCPF demonstrated the highest accuracy, no matter compared to SMOTE or NearMiss. Take CDC Diabetes. Here, ASCPF got to 94.12% in accuracy, pulled 87.99% sensitivity, hit 94.10% ROC-AUC - way ahead of SMOTE’s 77.65% accuracy and NearMiss’s 90.47%. It was important to keep the original proportions of groups to avoid to create misbalanced distribution and without changing the results it helped to achieve adding or removing the frequencies for cases of very few occurrences. One distinctive feature is that it is a combination of creating smart fake samples and ratio control – a novel approach which makes ASCPF unique. The combination results in more uniform performance on classes with skewed distribution, especially in important health decisions.

Read PDF

Similar papers

Open access 2026

A Comparative Analysis of Machine Learning Algorithms for the Early Prediction of Diabetes with an Evaluation of Class-Imbalance Handling

The study comes to the conclusion that headline accuracy is an unreliable guide in imbalanced medical prediction, that imbalance handling can change a model's practical usefulness, and that this benefit is strongly algorithm-dependent, meaning that the decision to resample should be based on the algorithm and the scree...

A. Oduroye, Temilade Opanuga, Esther Tosin Akanbi et al. · 0 citations
Open access Aug 2026

Enhancing Multiclass Diabetes Prediction with SMOTE and Machine Learning Classifiers

The results prove that ensemble models applied to class balancing and explainability methods can be used to rely on and provide clear tools regarding multiclass diabetes classification and illustrate the costs and benefits of operating a global model versus a minority-class sensitive one.

M. H. Ahmed, J. Qadir · 0 citations
Open access Sep 2026

Learning Under Extreme Class Imbalance: A Comparative Study of Algorithmic and Data-Level Solutions

A systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels indicates that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced perfo...

Tamsir Ariyadi, E. Noche, Nisha Pandey et al. · 0 citations
Open access Aug 2026

A Comparative Evaluation of Machine Learning Algorithms for Diabetes Risk Prediction

Evaluated machine learning algorithms for predicting diabetes risk from routinely available clinical and lifestyle variables confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained setti...

T. Olayinka · 0 citations

Cancer Risk Prediction Based on Integrated Machine Learning with Optimal Threshold Selection

The study suggests that this framework can serve as a clinical screening tool, prioritizing imaging examinations for high-risk individuals and using liquid biopsy for re-screening of moderate-risk groups, and driving the translation of precision cancer prevention from theory to practice.

Wei-Pei Liu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.