Skip to content
Open access

WHEN FAIR AI BECOMES UNFAIR: A COUNTERFACTUAL AUDIT OF POSITIONAL BIAS IN LARGE LANGUAGE MODELS FOR HIRING DECISIONS

Jul 2026 · Revista de Geopolítica · 0 citations · 29 references

TL;DR

Findings indicate that state-of-the-art LLMs can achieve a high degree of demographic neutrality; fundamental artefacts such as positional bias can nonetheless produce severely discriminatory outcomes; and bias auditing must extend beyond demographic parity to interaction artefacts and ecosystem structure.

Abstract

As Large Language Models (LLMs) are increasingly integrated into high-stakes recruitment processes, rigorous audits to detect algorithmic bias have become critical. This study implements a multi-phase auditing protocol to evaluate gender, racial/ethnic, and positional bias in state-of-the-art proprietary models (gpt-4.1-mini, GPT-5.2, Claude Sonnet 4.6, Gemini 2.5 Pro) and in an exploratory lower-capacity model. Using a simulated CEO-selection scenario with 60 functionally identical profiles, we employ counterfactual testing to isolate the effects of demographic attributes and presentation order. The results reveal a sharp divergence in model behaviour. Frontier models demonstrated remarkable neutrality, with no statistically significant evidence of gender or racial bias. In contrast, the exploratory model exhibited extreme segregation, ranking all female candidates above all male candidates in the original condition (Cliff's δ = 1.0, p < .001). Counterfactual analysis revealed that the root cause was not gender bias per se but an overwhelming primacy bias, whereby candidates presented at the beginning of the prompt were disproportionately favoured. Claude Sonnet 4.6 additionally displayed a form of "conscientious objection", declining to differentiate identical profiles. We complement the experimental audit with an ecosystem-level analysis of 2,653 provider-model listings from the open Models.dev catalogue, which corroborates the open-weights contamination-surface hypothesis but rejects the assumption that systemic monoculture exposure is confined to open-source models. These findings indicate that (a) state-of-the-art LLMs can achieve a high degree of demographic neutrality; (b) fundamental artefacts such as positional bias can nonetheless produce severely discriminatory outcomes; and (c) bias auditing must extend beyond demographic parity to interaction artefacts and ecosystem structure.

Read PDF

Similar papers

Preprint Jul 2026

FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation

Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants'names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.

Martin Lukk · 0 citations
Review Aug 2026

Counterfactual, Per-Decision Bias Auditing for Automated Hiring: Localizing and Explaining Disparate Impact in Applicant Tracking Systems

Automated applicant tracking systems increasingly decide who advances in hiring, and litigation and regulation now demand that those decisions be auditable. Existing tools sit at two extremes. Group fairness metrics such as the disparate impact ratio summarize a whole population but cannot say which individual decisions were unfair or why, while local explainers such as SHAP attribute a single prediction but are not connected to the legal standard by which hiring bias is judged. We present the AI Bias Firewall (AIBF), a method that audits an applicant tracking system one decision at a time. AIBF neutralizes a candidate's protected-attribute proxies, re-scores the decision, and measures the resulting counterfactual shift, which yields a signed per-decision bias in score points, a flag for decisions the protected attributes changed, and a plain-language explanation naming the responsible factors. We evaluate on two real public datasets, Adult and COMPAS, rather than on synthetic data. The per-decision counterfactual shift is faithful, aggregating to reproduce the known group level disparity, for example a mean shift of +7.5 points for the privileged group and -8.0 for the disadvantaged group on Adult, consistent with the measured statistical parity difference. AIBF identifies the decisions that protected attributes flipped with an area under the ROC curve of 0.963 on Adult, against 0.672 for a baseline that flags by group membership, and it identifies the harmed candidates so precisely that reviewing only five percent of decisions surfaces fifty-five percent of them, against six percent under group based review. We also report a limitation: correcting flagged decisions raises the disparate impact ratio substantially but not to legal parity, because features labeled as merit carry residual proxy correlation. AIBF is released under the Apache 2.0 license with code and experiments.

Jay Barach · 0 citations
Review Jul 2026

Analyzing and Correcting Benevolence Bias in Large Language Models

Benevolence bias is identified and measure, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions, and is easy to diagnose and straightforward to fix.

Yuanzi Li, Jun-Hao Wang, Minghui Liu et al. · 0 citations
Open access Jul 2026

Measuring Bias In Large Language Models: A Comparative Evaluation of LLaMA3.2-1B and LLaMA3.1-8B Across Indian Socio-Culture Dimension

In this paper we present a 30, 000-variant India-context bias audit to compare LLaMA3.2-1B and LLaMA3.1-8B across categories of Gender, Religion, Profession and Region that proposes the India Context Sensitivity Index (ICSI) as a category-weighted fairness metric. The larger model shows an improvement of the aggregate bias score that is statistically significant (Mann-Whitney p < 0.001) a biased-response rate that is significantly lower (χ² = 51.92, p < 0.001, 94.04% unbiased responses) at approximately double the inference latency. The breakdown of results by categories also shows that this overall gain is unevenly split: the 8B remains more sensitive on Gender-based prompts even as it improves on Religion, Profession and Region, underscoring the merit of reporting disaggregated, category-imbued fairness over a single number bias score for India deployed LLMs [9], [17], [23]. The future work will allow the scaling of the benchmark to be done to additional model sizes within the LLaMA family (for instance 3B, 70B) to instead fit a bias-versus-scale curve rather than a two-point comparison (14). Then, the extension of the category set to ‘caste-adjacent’ and intersectional categories (for instance gender × region) that are under-represented in this effort (18), (9). The next direction will be replacing the lexicon-based bias scorer with the LLM-as-judge scorer validated against human annotation, to measure scorer-induced bias in the evaluation pipeline itself (28). Finally, deploying the mitigation variants (few-shot, prompt-engineering) evaluated here as live runtime guardrails and measuring their effect on the ICSI in a closed deployment loop (29), (30).

Neha Neha, Amandeep Noliya · 0 citations
Conference Open access 2026

Bias and Fairness in LLM-Based Recruitment: A Systematic Review

As large language models emerge as critical infrastructure in labor markets, questions about their governance touch on some of the most contested issues at the intersection of AI development, complex sociotechnical systems and emerging technology regulation. Recruitment tools powered by these models are now widely deployed across industries and they raise urgent concerns about intersectional algorithmic discrimination—patterns of unequal treatment that remain invisible so long as auditors examine only one protected attribute at a time. Trained on historical hiring data, these systems tend to reproduce and in some cases deepen, discriminatory patterns that cut across gender, race and other protected characteristics simultaneously. Two research communities bear on this problem without yet speaking to each other adequately: the fairness-in-NLP literature and the human-centered AI (HCAI) governance literature have each grown substantially, but largely in parallel. To examine where they diverge and why, we conducted a PRISMA 2020-guided systematic literature review drawing on 82 studies selected from 493 records retrieved from Scopus and Web of Science. Grounded theory coding and a concept matrix organize the findings around four thematic axes: bias sources across the hiring pipeline, formal fairness metrics and their mathematical limits, debiasing techniques and governance frameworks. What the concept matrix reveals is, above all, a structural disconnect. Technical studies rarely engage with HCAI design principles governance-oriented work rarely operationalizes the technical limitations that the empirical literature has documented in detail. Five evidence-based design recommendations follow from the analysis; a particularly urgent recommendation concerns the development of non-Western fairness benchmarks. A targeted research agenda addresses intersectional auditing and LLM-specific debiasing as the two highest priority open problems.

Asmae El Moutafail, Khalid Belkhoutout · 0 citations
Open access Aug 2026

condfair: An R Package for Ability-Conditioned Fairness and Explanation Diagnostics in Automated Scoring.

Fairness in automated scoring is typically evaluated with a single global statistic contrasting a focal and reference group-an approach that can either mask a disparity that changes sign across the ability range, or overstate one by conflating it with genuine ability differences between groups (impact). We introduce condfair, an R package that adapts differential item functioning (DIF) logic to automated scoring: it estimates a conditional disparity function across ability levels, tests it with a wild-bootstrap omnibus procedure, decomposes bias into uniform and non-uniform components, and identifies candidate feature-level sources of a detected disparity via conditional SHAP disparity testing. Using the PERSUADE 2.0 essay corpus, we show a global measure can conceal a large, ability-concentrated gender disparity (marginal gap = 0.003; peak conditional disparity = 0.276, p = .001) while overstating an English Language Learner disparity by conflating it with impact (marginal gap = 0.280; conditional bias = 0.045).

Tri Zahra Ningsih, Aman Aman, Ahmad Nasrulloh · 0 citations