Skip to content
Open access

Learning Under Extreme Class Imbalance: A Comparative Study of Algorithmic and Data-Level Solutions

Sep 2026 · Journal of Data Science · 0 citations · 20 references

TL;DR

A systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels indicates that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced performance across all scenarios.

Abstract

Extreme class imbalance remains a persistent challenge in machine learning, particularly in high-impact domains such as fraud detection, medical diagnosis, and risk analysis, where minority classes represent critical outcomes. Conventional models often fail in such settings due to their bias toward majority classes, resulting in poor minority detection despite high overall accuracy. Although various data-level and algorithm-level techniques have been proposed, existing studies typically evaluate them in isolation and lack a comprehensive understanding of their effectiveness across different imbalance conditions. To address this gap, this study proposes a systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels. Multiple benchmark datasets are utilized, and experiments are conducted using standardized preprocessing, controlled imbalance simulation, and repeated trials to ensure robustness. Performance is assessed using imbalance-aware metrics, including precision, recall, F1-score, and ROC-AUC. The results indicate that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced performance across all scenarios. In particular, hybrid models demonstrate superior minority class recall and F1-score while maintaining competitive precision, especially under extreme imbalance conditions. The primary goal of this research is to provide a comprehensive evaluation of imbalance-handling strategies and offer practical guidance for selecting appropriate techniques based on dataset characteristics. The findings highlight the importance of combining data-centric and model-centric approaches to enhance robustness and reliability in imbalanced learning environments. The results demonstrate up to a 32% improvement in recall compared to baseline models.

Read PDF

Similar papers

Open access Sep 2026

A Comparative Performance Evaluation of Classification Algorithms on Imbalanced Datasets

Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rar...

Necati Vardar, Mehmet Fatih Ören · 0 citations
Open access Sep 2026

Robust Evaluation Metrics for Assessing Machine Learning Performance Beyond Accuracy

A structured multi-metric evaluation framework that integrates classification, ranking-based, calibration, and robustness-oriented metrics for comprehensive ML performance assessment is proposed capable of supporting more reliable real-world ML deployment.

Ade Putra, E. Noche, Diksha D. Gabhane · 0 citations
Open access Sep 2026

Adaptive Weighting–Synthetic Minority Oversampling Technique

A novel oversampling algorithm: the adaptive weighting–synthetic minority oversampling technique (AW-SMOTE), which combines the two perspectives of boundary tightness and local density and provides global sample enhancement support.

Shen Yan, Hai-Feng Guo, Xiao-Ming Su · 0 citations
Open access 2026

Optimizing diabetes prediction in machine learning models: Evaluating the effectiveness of a novel class imbalance technique—adaptive synthetic class balancing with class proportion filtering

Adaptive Synthetic Class Balancing with filtering by class ratio (ASCPF) is suggested, a combination of creating smart fake samples and ratio control which results in more uniform performance on classes with skewed distribution, especially in important health decisions.

Pankaj Beldar, Snehal M. Kamalapur, Priti Vaidya et al. · 0 citations
Open access Sep 2026

Enhancing Imbalanced Data Classification with Class-Aware Synthetic Minority Over-sampling Technique

Imbalanced data classification is one of the difficult machine learning tasks in which the majority class outnumbers the minority classes. This problem is prevalent in various domains such as medical disease detection, spam/fraud detection, digital marketing, agriculture and telecommunications. Synthetic Minority Over-...

Ruturaj Mahajan, S. Patil, Vilabha Patil · 0 citations
Open access Aug 2026

Performance Evaluation of Classical Machine Learning Models for Insurance Fraud Detection Under Severe Class Imbalance

Experimental results demonstrate that ensemble-based methods, particularly Random Forest, achieve superior performance in identifying minority class fraud cases while maintaining stable overall classification accuracy in fraud detection scenarios.

Garima Sharma · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.