Sep 2026· Journal of Data Science· 0 citations· 24 references
TL;DR
A structured multi-metric evaluation framework that integrates classification, ranking-based, calibration, and robustness-oriented metrics for comprehensive ML performance assessment is proposed capable of supporting more reliable real-world ML deployment.
Abstract
The widespread deployment of machine learning (ML) systems in critical domains has exposed the limitations of accuracy-centric evaluation, particularly under conditions involving class imbalance, noise, and distributional shifts. Existing studies frequently employ alternative metrics in isolation and lack a unified framework capable of systematically assessing model robustness, reliability, and decision sensitivity across varying data conditions. To address this gap, this study proposes a structured multi-metric evaluation framework that integrates classification, ranking-based, calibration, and robustness-oriented metrics for comprehensive ML performance assessment. A quantitative experimental design is employed using multiple benchmark datasets with varying statistical characteristics, including balanced and imbalanced distributions. Controlled perturbation scenarios—including noise injection, class imbalance manipulation, and distribution shift simulation—are introduced to emulate realistic deployment environments. Several machine learning models, namely Logistic Regression, Support Vector Machines, Random Forest, and Multi-Layer Perceptron (MLP), are evaluated using metrics such as Accuracy, F1-score, ROC-AUC, PR-AUC, Brier Score, and Expected Calibration Error (ECE). The experimental results demonstrate that accuracy consistently overestimates model effectiveness under adverse conditions, while alternative metrics reveal substantial hidden weaknesses in minority class detection and probability reliability. Among the evaluated models, MLP achieved the strongest overall performance, obtaining a ROC-AUC of 0.94 and PR-AUC of 0.89 under baseline conditions. Furthermore, calibration-oriented metrics exhibited significantly higher sensitivity to perturbation severity compared to accuracy. This study contributes to the advancement of trustworthy artificial intelligence by promoting a comprehensive, context-aware, and robustness-oriented evaluation framework capable of supporting more reliable real-world ML deployment.
Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rar...
Necati Vardar, Mehmet Fatih Ören· Sakarya University Journal o...· 0 citations
A systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels indicates that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced perfo...
Tamsir Ariyadi, E. Noche, Nisha Pandey et al.· Journal of Data Science· 0 citations
Machine learning research has traditionally emphasized model-centric optimization, often overlooking the critical role of data quality in determining generalization performance. However, real-world datasets frequently suffer from noise, imbalance, and limited diversity, which constrain model effectiveness despite incre...
Nia Oktaviani, E. Noche, S. Patil et al.· Journal of Data Science· 0 citations
Tree-based models exhibited the largest reductions in accuracy and the steepest increases in calibration error under scarcity, whereas logistic regression preserved threshold balance and probability calibration at negligible computational cost.
Jairo Devon A. Daquioag, E. Ayo· Engineering and Technology J...· 0 citations
Logistic Regression, Random Forest, soft voting, and weighted voting using a publicly available educational dataset comprising 4,424 student records, 36 predictors, and three outcome classes provided accurate, interpretable, and decision-oriented predictions, although external validation and prospective evaluation are...
Haleeful Jud· Journal of Intelligent Decis...· 0 citations
A Standardized Multi-Metric Evaluation Framework is developed to support the selection of more objective and reproducible machine learning models in the case of loan approval to identify models that have a balance of performance and efficiency in loan approval experiments.
Trihartono Agus, Agus Ilyas Ilyas, S. Sattriedi et al.· Jurnal Informatika: Jurnal P...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.