Skip to content
Open access

A Quantitative and Qualitative Analysis of Data Selection Impact on Machine Learning Fairness and Utility

2026 · IEEE Access · Vol 14, pp. 130881-130901 · 1 citation · 75 references

Abstract

With the continuous growth in data volumes, efficient data selection has become critical for scaling machine learning (ML) classifiers. Although the existing methods aim to accelerate training without compromising accuracy, their impact on fairness and broader utility remains largely unexplored. This paper presents the first quantitative and qualitative evaluation of these effects through a large-scale empirical analysis comprising 24,395 experiments across 9 real-world datasets, 7 ML models, and several state-of-the-art data selection methods, conducted using ELAPSE, our open-source evaluation framework. Our analysis reveals several key findings and non-trivial observations. The results show that ML data selection can hurt model fairness in a non-negligible number of cases, and compromise model utility in more than half of the cases. These effects differ across datasets, with no single method consistently outperforming the others, and in nearly one-third of the cases, either utility or fairness must be prioritized. Surprisingly, combining data selection with classical bias mitigation increases model unfairness and often degrades model utility. These findings provide interesting research directions for utility- and fairness-aware ML data selection.

Read PDF