Skip to content

Beyond the Name: Demographic Leakage in De-Identified R\'esum\'es and Evaluation Artifacts in LLM Bias Audits

Sep 2026 · 0 citations · 16 references
Computer Science

Abstract

De-identified r\'esum\'e screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-weight models and 620 counterfactual r\'esum\'es. By holding language attributes strictly identical, we isolate unstructured prose across five ethnocultural conditions and three cue-salience tiers. Target-group recovery averages 0.757 overall and saturates at 1.000 under high salience, demonstrating that non-language prose sustains demographic inference. Crucially, models diverge only under faint cues (0.086-0.690), establishing salience as an essential evaluation axis. Furthermore, pairwise LLM-as-a-judge outcomes are highly sensitive to evaluation design: forbidding ties yields an apparent selection-rate ratio of 0.39 alongside strong position and content effects, whereas permitting ties produces near-universal ties for most models ($\ge94\%$). Downstream scoring shows only very small between-condition differences, highlighting the need to distinguish demographic signals recoverable from r\'esum\'e content from effects introduced by the evaluation protocol.

View source

Similar papers

Preprint Aug 2026

It's How You Ask: Gender-Associated Linguistic Bias in LLMs

Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses ac...

Katherine Van Koevering, Anjalie Field · 0 citations
#artificial intelligence Preprint Sep 2026

Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment

Open-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly understood. We present the first systematic, multi-m...

Kosuke Kitahara, Nobuhiro Yamaguchi · 0 citations
Conference Open access 2026

Bias and Fairness in LLM-Based Recruitment: A Systematic Review

A PRISMA 2020-guided systematic literature review draws on 82 studies selected from 493 records retrieved from Scopus and Web of Science and reveals a structural disconnect in the fairness-in-NLP and HCAI governance literature.

Asmae El Moutafail, Khalid Belkhoutout · 0 citations
Open access Aug 2026

“This is how journalism has always been done”: White neutrality, newsroom norms, and an AI intervention in source diversification

This study draws on interviews with 12 journalists in a Canadian newsroom to examine how they understand source diversification, how they perceive AI-based source-tracking technology, and how their reflections illuminate discourse on race and accountability in journalism. Based on semi-structured interviews conducted a...

Gavin Adamson, A. Malik, Shari Okeke et al. · 0 citations
#natural language process... Preprint Sep 2026

Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models

A robust Political Compass Test evaluation framework is introduced that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording, and results show most models lean Libertarian-Left on average and larger models sho...

Luka Debevc, Nishan Chatterjee, Antoine Doucet et al. · 0 citations
Open access Sep 2026

“Unprecedented Injustice”: network determinism in the Dutch digital welfare state

Automated decision-making systems in the welfare state have produced discriminatory outcomes at scale, as critical scholarship has extensively documented. This article examines the mechanisms through which such outcomes are produced by tracing the cultural, institutional, and technical determinants embedded in their de...

Diletta Huyskes · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.