Skip to content
Open access

Robust Evaluation Metrics for Assessing Machine Learning Performance Beyond Accuracy

Sep 2026 · Journal of Data Science · 0 citations · 24 references

TL;DR

A structured multi-metric evaluation framework that integrates classification, ranking-based, calibration, and robustness-oriented metrics for comprehensive ML performance assessment is proposed capable of supporting more reliable real-world ML deployment.

Abstract

The widespread deployment of machine learning (ML) systems in critical domains has exposed the limitations of accuracy-centric evaluation, particularly under conditions involving class imbalance, noise, and distributional shifts. Existing studies frequently employ alternative metrics in isolation and lack a unified framework capable of systematically assessing model robustness, reliability, and decision sensitivity across varying data conditions. To address this gap, this study proposes a structured multi-metric evaluation framework that integrates classification, ranking-based, calibration, and robustness-oriented metrics for comprehensive ML performance assessment. A quantitative experimental design is employed using multiple benchmark datasets with varying statistical characteristics, including balanced and imbalanced distributions. Controlled perturbation scenarios—including noise injection, class imbalance manipulation, and distribution shift simulation—are introduced to emulate realistic deployment environments. Several machine learning models, namely Logistic Regression, Support Vector Machines, Random Forest, and Multi-Layer Perceptron (MLP), are evaluated using metrics such as Accuracy, F1-score, ROC-AUC, PR-AUC, Brier Score, and Expected Calibration Error (ECE). The experimental results demonstrate that accuracy consistently overestimates model effectiveness under adverse conditions, while alternative metrics reveal substantial hidden weaknesses in minority class detection and probability reliability. Among the evaluated models, MLP achieved the strongest overall performance, obtaining a ROC-AUC of 0.94 and PR-AUC of 0.89 under baseline conditions. Furthermore, calibration-oriented metrics exhibited significantly higher sensitivity to perturbation severity compared to accuracy. This study contributes to the advancement of trustworthy artificial intelligence by promoting a comprehensive, context-aware, and robustness-oriented evaluation framework capable of supporting more reliable real-world ML deployment.

Read PDF

Similar papers

Open access Sep 2026

A Comparative Performance Evaluation of Classification Algorithms on Imbalanced Datasets

Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rar...

Necati Vardar, Mehmet Fatih Ören · 0 citations
Open access Sep 2026

Learning Under Extreme Class Imbalance: A Comparative Study of Algorithmic and Data-Level Solutions

A systematic comparative framework that integrates data-level resampling, cost-sensitive learning, and hybrid approaches to evaluate their performance under varying imbalance ratios and noise levels indicates that hybrid approaches consistently outperform standalone methods, achieving the most stable and balanced perfo...

Tamsir Ariyadi, E. Noche, Nisha Pandey et al. · 0 citations
Open access Sep 2026

Evaluating Data-Centric Optimization Strategies for Improving Machine Learning Generalization

Machine learning research has traditionally emphasized model-centric optimization, often overlooking the critical role of data quality in determining generalization performance. However, real-world datasets frequently suffer from noise, imbalance, and limited diversity, which constrain model effectiveness despite incre...

Nia Oktaviani, E. Noche, S. Patil et al. · 0 citations
#software testing Open access Nov 2026

How Classification Models Degrade as Data Diminish: A Multi-Criteria Analysis of Tabular Classifiers Under Controlled Data Scarcity

Tree-based models exhibited the largest reductions in accuracy and the steepest increases in calibration error under scarcity, whereas logistic regression preserved threshold balance and probability calibration at negligible computational cost.

Jairo Devon A. Daquioag, E. Ayo · 0 citations
Review Open access Aug 2026

Hybrid Machine Learning Models for Predictive Analytics and Intelligent Decision-Making in Complex Data Environments

Logistic Regression, Random Forest, soft voting, and weighted voting using a publicly available educational dataset comprising 4,424 student records, 36 predictors, and three outcome classes provided accurate, interpretable, and decision-oriented predictions, although external validation and prospective evaluation are...

Haleeful Jud · 0 citations
Open access Aug 2026

Toward Reliable Machine Learning Model Selection: A Standardized Multi-Metric Evaluation Framework For Loan Approval

A Standardized Multi-Metric Evaluation Framework is developed to support the selection of more objective and reproducible machine learning models in the case of loan approval to identify models that have a balance of performance and efficiency in loan approval experiments.

Trihartono Agus, Agus Ilyas Ilyas, S. Sattriedi et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.