A Systematic Review of Bias Detection and Dataset Improvement Methods for Fair AI Systems
Algorithmic fairness refers to the equitable treatment of users regardless of their identity, demographics, or other personal characteristics, whereas biased AI may produce unfair outcomes when the input data reflect imbalances or underrepresentation of certain groups. This systematic literature review analyzes studies published between 2018 and 2025 that report approaches for handling bias in data and datasets, with particular attention to dataset-centered techniques. We examine how bias is identified, mitigated, and reflected in model outcomes. Our method defines research questions and details the search strategy, study selection procedure, and protocols for data extraction and analysis. The review maps fairness dimensions, the metrics employed, pre-processing and in-processing methods, the role of human participation, observed effects and side effects, and the ethical considerations that inform practice. The evidence suggests that proactive, iterative approaches incorporating human oversight are essential for mitigating algorithmic discrimination and enhancing the trustworthiness of AI systems. Neglecting these considerations risks significant societal harm and the perpetuation of bias against marginalized groups in high-stakes contexts, particularly in an era where AI is increasingly widespread.