Skip to content

Moving Towards a Fairer Future: Reproducing Kamiran and Calders’ Findings on Anti-Discrimination Techniques

Unknown authors
· 0 citations · 23 references

TL;DR

A reproduces three of Kamiran and Calders' empirical findings and extends the work by investigating how ranker choice affects the performance of Massaging and Preferential Sampling, recommending Massaging as the most reliable technique for practical deployment.

View source

Similar papers

#software testing Open access Aug 2026

FIFT: Feature Importance-Guided Fairness Testing for machine learning software

Results indicate that global feature importance, used as an active search signal rather than a post-hoc diagnostic, improves both the effectiveness and the efficiency of individual fairness testing.

H. Mamman, Abdullateef Oluwagbemiga Balogun, Mustapha Maidawa et al. · 0 citations
Preprint Aug 2026

The Overstated Cost of AI Fairness in Criminal Justice

A dominant critique of algorithmic fairness holds that increasing fairness reduces predictive accuracy, imposing a cost on society. We challenge that assumption by empirically analyzing the COMPAS dataset. We make two contributions. First, using causal inference methods, we show that racial bias is not only present in the COMPAS dataset but is also amplified by the models trained on it. Widely used models do more than replicate existing bias; they exacerbate it. This undercuts both the assumption that algorithmic decision-making offers a neutral improvement over human judgment and the weaker claim that it merely mirrors preexisting human bias. Second, we reframe the fairness-accuracy tradeoff. Applying fairness constraints does not necessarily cost predictive accuracy in criminal justice. Prediction systems operationalize concepts such as risk through implicit and often flawed normative choices about what to predict and how. The tradeoff claim assumes that the unconstrained model's prediction is an optimal baseline. Fairness constraints can instead correct distortions introduced by biased outcome variables: rearrest data, in this case, captures and magnifies systemic racial disparities. Under some interventions, therefore, fairness carries none of the cost presumed in policy debates. These dynamics extend beyond criminal justice to lending, hiring, and housing, where biased outcome variables reinforce inequality independently of proxy selection. We draw out what this implies for how law and policy should approach fairness adjustments in criminal law.

Ignacio Cofone, Warut Khern-am-nuai University of Oxford, McGill University · 3 citations
Open access Aug 2026

Bounding the Fairness of a Classifier Using Population-level Statistics

This work introduces a method to lower-bound the discrepancy of a classifier: a quantity that jointly captures inaccuracy and unfairness, and develops a computationally efficient procedure for calculating the tightest possible lower bound on the classifier’s discrepancy.

Sivan Sabato, E. Yom-Tov · 0 citations
Preprint Jul 2026

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants'names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.

Martin Lukk · 0 citations
Open access Jul 2026

WHEN FAIR AI BECOMES UNFAIR: A COUNTERFACTUAL AUDIT OF POSITIONAL BIAS IN LARGE LANGUAGE MODELS FOR HIRING DECISIONS

Findings indicate that state-of-the-art LLMs can achieve a high degree of demographic neutrality; fundamental artefacts such as positional bias can nonetheless produce severely discriminatory outcomes; and bias auditing must extend beyond demographic parity to interaction artefacts and ecosystem structure.

A. Camargo, Rafaela Silva Figueiredo Camargo · 0 citations
Open access 2026

A Multi-Stage Framework for Bias Detection and Mitigation in AI-Driven Recruitment Systems

The use of machine learning in recruitment has raised growing concerns about fairness, as automated hiring systems can generate unequal outcomes across demographic groups. These disparities are influenced not only by imbalanced data but also by the behavior of learning algorithms, making bias a multidimensional challenge that cannot be effectively addressed with single-stage solutions. This study introduces an integrated framework for bias detection and mitigation in AI-driven recruitment systems, combining interventions at the data, model, and decision levels within a unified evaluation pipeline. The framework is assessed using multiple classification models of varying complexity and evaluated with established fairness metrics. In addition, explainability techniques are employed using SHAP-based feature attribution to investigate hidden dependencies and assess the sensitivity of predictions to demographic attributes. Experimental results show that baseline models achieve strong predictive performance, with accuracy ranging from 0.807 to 0.816; however, fairness evaluation reveals substantial disparities, with Disparate Impact as low as 0.190 and Demographic Parity Difference exceeding 0.27 in some cases. After applying the proposed multi-stage mitigation approach, fairness metrics improve significantly: Disparate Impact meets or exceeds the legal threshold of 0.80 across all models, with reductions in Demographic Parity Difference of 70–85% and Equal Opportunity Difference of 47–75%, and demographic disparities are reduced to 0.029–0.053. These improvements are achieved with minimal performance trade-off, as overall accuracy decreases by at most 4.2 % points while ROC AUC remains unchanged. The findings demonstrate that bias in recruitment systems arises from the interplay between data and model dynamics and highlight the importance of coordinated mitigation strategies throughout the machine learning lifecycle. This work provides a practical, scalable approach to developing fair and transparent AI systems for hiring applications.

Gideon Assafuah, Claude Turner, C. Turner et al. · 0 citations