Optimizing diabetes prediction in machine learning models: Evaluating the effectiveness of a novel class imbalance technique—adaptive synthetic class balancing with class proportion filtering
2026· Sigma Journal of Engineering and Natural Sciences· 0 citations· 35 references
TL;DR
Adaptive Synthetic Class Balancing with filtering by class ratio (ASCPF) is suggested, a combination of creating smart fake samples and ratio control which results in more uniform performance on classes with skewed distribution, especially in important health decisions.
Abstract
The skewedness of results when predicting diabetes is mostly due to uneven distribution of data, especially in reducing detection rates of real patients. These are errors which cause delay in treatment or incorrect diagnosis. This work suggests a counter plan to this assumption, which is Adaptive Synthetic Class Balancing with filtering by class ratio (ASCPF). This does not require siloing such methods as SMOTE or NearMiss, but it does change the way to create or make samples, as well as the way to achieve retention, depending on group size. We evaluate the performance of each strategy by considering rare cases, using Extra Trees Classifier. We ran tests on the following data collections: CDC Diabetes, Breast Cancer Wisconsin (Diagnostic), and Credit Card Fraud Detection, KDD Cup 1999 Intrusion Detection, PIMA, and BRFSS. In the performance trend, ASCPF demonstrated the highest accuracy, no matter compared to SMOTE or NearMiss. Take CDC Diabetes. Here, ASCPF got to 94.12% in accuracy, pulled 87.99% sensitivity, hit 94.10% ROC-AUC - way ahead of SMOTE’s 77.65% accuracy and NearMiss’s 90.47%. It was important to keep the original proportions of groups to avoid to create misbalanced distribution and without changing the results it helped to achieve adding or removing the frequencies for cases of very few occurrences. One distinctive feature is that it is a combination of creating smart fake samples and ratio control – a novel approach which makes ASCPF unique. The combination results in more uniform performance on classes with skewed distribution, especially in important health decisions.
The study comes to the conclusion that headline accuracy is an unreliable guide in imbalanced medical prediction, that imbalance handling can change a model's practical usefulness, and that this benefit is strongly algorithm-dependent, meaning that the decision to resample should be based on the algorithm and the scree...
A. Oduroye, Temilade Opanuga, Esther Tosin Akanbi et al.· International journal of re...· 0 citations
The results prove that ensemble models applied to class balancing and explainability methods can be used to rely on and provide clear tools regarding multiclass diabetes classification and illustrate the costs and benefits of operating a global model versus a minority-class sensitive one.
M. H. Ahmed, J. Qadir· Academic Journal of Internat...· 0 citations
A systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels indicates that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced perfo...
Tamsir Ariyadi, E. Noche, Nisha Pandey et al.· Journal of Data Science· 0 citations
The pipeline approach ensured no data leakage in cross-validation, and the findings support ensemble ML models with SMOTE as a preprocessing step for imbalanced CVD datasets.
M. Maindarkar· Journal of Intelligent Decis...· 0 citations
Evaluated machine learning algorithms for predicting diabetes risk from routinely available clinical and lifestyle variables confirm that ensemble tree-based methods, particularly Random Forest, provide a reliable, interpretable, and deployable basis for diabetes risk screening, especially in resource-constrained setti...
T. Olayinka· FUDMA Journal of Sciences· 0 citations
The study suggests that this framework can serve as a clinical screening tool, prioritizing imaging examinations for high-risk individuals and using liquid biopsy for re-screening of moderate-risk groups, and driving the translation of precision cancer prevention from theory to practice.
Wei-Pei Liu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.