A side effect that misrepresents patients is measured: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes the question never stated, in effect rewriting who the patient is.
Abstract
Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes (race, socioeconomic status, sex) the question never stated, in effect rewriting who the patient is. We call this demographic injection. Across 47 models, four medical benchmarks, and 376,000 responses scored by a validated model-judge pipeline, a single DEI prompt raises the injection rate from 0.7% to 33.1% (47x) in all 47 of 47 models, attributable to the equity content rather than to added length (18x above a length-matched control; p=1.4x10^-14). Most added content is a general population statement that leaves the answer unchanged, but a smaller subset attaches an attribute to the specific patient or changes the selected option (0.25-2.4% of responses, 99.8% toward the incorrect option), where the invented demographic changes the answer the model recommends. Phrasing scales the effect from 14% to 56%. DEI prompts are just one example of a more general mechanism. Any instruction that nudges how a model reasons can make it add unrequested details, including details about the patient. Flagged outputs are treated as model errors under study, not clinical guidance.
Five widely used prompting strategies across five influential LLMs in the latest medical bias benchmark reveal substantial heterogeneity in both effectiveness and overhead across models, with no strategy proving universally effective and some even exacerbating bias.
Ying Xiao, Zhenpeng Chen, Jie M. Zhang· Philosophical transactions....· 1 citation
MedQAbstain is introduced, a benchmark explicitly designed to evaluate medical abstention under uncertainty, and finds that state-of-the-art LLMs systematically overcommit, rarely abstaining even when the question itself is hidden.
Alessio Cocchieri, Luca Ragazzi, Giuseppe Tagliavini et al.· Annual Meeting of the Associ...· 2 citations
VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.
Praveen Reddy, C. Mandke, Suvrankar Datta et al.· 0 citations
Whether clinical safety established in English transfers to Hausa is asked, and whether any failure is attributable to the language, the clinical task, or the class of model that low-resource deployment admits.
Anthonio Oladimeji Gabriel, Dimeji AbdulSobur Olawuyi, T. Ajayi et al.· 0 citations
An LLM leaderboard showing how open-source LLMs perform at entity extraction on unseen clinical notes is developed, showing that large language models are already available that can perform entity extraction well enough to be considered in place of some administrative data.
E. Martin, Seungwon Lee, K. Riazi et al.· International Journal of Pop...· 0 citations
It is found that LLM access enhances performance on standardized clinical vignettes in all three countries, and policymakers should prioritize structured integration of LLMs as decision-support tools, combined with targeted training, local validation, and safeguards against automation bias rather than relying on access alone.
N. Rounding, L. S. Arif, Janine Berg et al.· 0 citations