A prespecified randomized algorithm audit of what causally moves large language model (LLM) assistants' recommendations, finding that gender and ethnicity were signaled through names following correspondence-audit methodology.
Abstract
Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models'own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.
Findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues, which position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.
A. Kapadia, Eshwar Chandrasekharan, Koustuv Saha· 0 citations
Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context examples raise it above zero in every model. Accuracy-based evaluation detects none of this.
How can you tell whether a defense attorney is any good? In this article, we describe a methodology for measuring the quality of legal representation by comparing machine learning predictions based on objective case characteristics with actual case outcomes. Our premise is that if an attorney's cases consistently end better than expected, they are probably doing their job well. Similar logic has long been used in other fields: it is how teachers are evaluated in the economics of education, surgeons in healthcare, and players in professional sports. We adapt these approaches to criminal proceedings. Such a predictive model can be trained on Russian court data from 2010–2025 and would account for the fact that a criminal case passes through several sequential procedural decision points. The resulting rating is adjusted for teamwork among multiple defense attorneys, the stage of proceedings, and small caseloads. We discuss the weaknesses of the approach: most importantly, the pretrial stage is invisible to us, published decisions are incomplete, and skilled attorneys may systematically select more difficult cases. Nevertheless, even in this form, such an evaluation tool can provide clients with a meaningful and objective benchmark when choosing a defense attorney, which could substantially improve the quality of information in the legal services market.
A. Kazun, V. Devyatnikov, Mikael Belov· Issues of Economic Theory· 0 citations
Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining three were imprecise or borderline, with no consistent elite advantage. Association strength did not reliably predict decision leakage across models (r = 0.201, p = 0.633) or across frozen surname-pair-by-model cells (r = 0.065, p = 0.565). The central result is a measurement dissociation: latent social association and consequential treatment are empirically distinct constructs. Evaluations should measure the transition from association to action directly.
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such committees can be gamed by shortcuts, cues a benchmark rewards but a clinician would ignore. Across seven cohorts on six public datasets spanning text (MedQA-USMLE, MedMCQA, MIMIC-CXR reports), imaging (NIH ChestX-ray14, MIMIC-CXR-JPG, CheXpert) and tabular ICU records (SUPPORT2), Gemini committees resist these cues in isolation (flip 5-16%), yet a socially plausible shortcut spreads: when two peers assert the same wrong answer, the holdout under test adopts it in 38% of cases, as does a false"pre-screen"system flag, on both capability tiers. Of three oversight agents, a gate cannot separate adoption from honest agreement (false-positive rate 100%); a same-lineage judge reading only the transcript flags adoption on text (precision 100%, recall 93%) but collapses onto the gate in imaging; a referee that privately re-queries the holdout transfers to imaging (77-88% precision, 13-21% false-positive rate). Tripling a cue's visual salience does not move contagion, whereas a second peer voice raises it by half again. Gaming a hidden rubric is near-silent: only 1/10 text and 1/134 imaging drifters name the rubric they moved toward. What games a committee is social plausibility, and only a referee independent of self-report catches it. Code: https://github.com/criticaldata/benchmaxxing
S. Ordóñez, Agastya Munnangi, Aldo Marzullo et al.· 0 citations
Background: Artificial intelligence (AI) systems, including large language models, are increasingly used in clinical practice, whether consulted informally by clinicians or introduced by employers into decision workflows. It remains unclear how physicians attribute moral responsibility when a decision follows an AI recommendation and whether that attribution varies with the type of decision at stake. Methods: We conducted a cross-sectional, within-subject vignette survey of physicians in Romania. Each respondent rated the same three scenarios—urgent clinical, elective clinical, and administrative—in which a physician followed an AI recommendation under two extenuating institutional constraints. Five-point Likert items addressed the mitigation of blame by circumstances, physician responsibility despite the AI recommendation, and institutional co-responsibility. Analyses were non-parametric, with corrections for multiple testing. Results: Among 72 physicians from 17 specialties, respondents endorsed full personal responsibility in every scenario, including the administrative one, with no significant difference between scenarios. They rejected extenuating circumstances as mitigating in both clinical scenarios but were divided about them in the administrative scenario, which had the largest effect. Institutional co-responsibility was endorsed alongside personal responsibility rather than in place of it, and the two attributions were largely uncorrelated. No demographic association survived correction, although the study was not powered to detect small-effect sizes. Conclusions: Physicians treated AI as an instrument rather than a bearer of responsibility, which is unsurprising. The substantive findings lie elsewhere: personal responsibility was retained across all decision contexts, while what varied was the admissibility of institutional constraints as excuses and the emphasis placed on the institution’s share. Because respondents did not treat responsibility as a fixed quantity to be divided, the pattern is consistent with distributed-responsibility accounts rather than with a responsibility gap, though attitudinal data cannot adjudicate between normative accounts. The findings are exploratory and require confirmation in larger, more representative samples.
Florian Berghea, Alexandra-Ligia Dincă, D. Ciuc et al.· Applied Sciences· 0 citations