Skip to content
Open access

Profile-associated financial and access-related framing in LLM-generated pediatric asthma referral plans: a factorial audit of seven large language models

Jul 2026 · Frontiers in Digital Health · Vol 8 · 0 citations · 25 references
Medicine

TL;DR

LLM-generated pediatric asthma referral plans varied in financial-access, geographic-access, navigation, SDOH-recognition, and selected tone-related framing, which support evaluating structural and access-related framing alongside biomedical content in clinical LLM audits.

Abstract

Background Large language models (LLMs) are increasingly considered for clinical documentation, referral support, and patient-facing communication. Biomedical accuracy alone may be insufficient for safe deployment if generated plans vary in access navigation, financial-access language, referral specificity, or tone across socially meaningful patient cues. Objective To evaluate whether onomastic and bundled geographic-access signals are associated with differences in LLM-generated pediatric asthma referral plans. Methods We conducted a cross-sectional 2 × 2 factorial audit of seven commercial LLMs. A standardized vignette described a 5-year-old boy with moderate persistent asthma, persistent nocturnal symptoms, FEV1 of 70% predicted, and an Asthma Control Test score of 16. Patient name, Liam Miller versus DeShawn Washington, and address/geography, Palo Alto, CA versus Indianola, MS, were manipulated while clinical facts were held constant. Each model generated 20 responses per profile, yielding 560 referral plans. Outputs were scored using a prespecified Automated Structural Competence Scoring framework. The primary endpoint was response-length-adjusted M16 Financial-Access Term Rate, analyzed using a negative-binomial model with log word-count offset and LLM fixed effects. Key secondary endpoints were controlled using Benjamini-Hochberg false-discovery-rate correction. Results In the response-level primary model, the DeShawn name signal was associated with a higher financial-access term rate (IRR, 1.47; 95% CI, 1.23–1.77; p < 0.001), as was the bundled geographic-access signal (IRR, 2.40; 95% CI, 2.02–2.85; p < 0.001). The name-signal association was directionally similar but less precise in model-profile aggregated sensitivity analysis. The interaction term was below 1.0 (IRR, 0.79; 95% CI, 0.63–1.00; p = 0.048), indicating no positive multiplicative synergy. Institutional Specificity and Triage Ranking were at ceiling. SDOH Recognition Depth, Location-Friction Acknowledgment, Navigator Recommendation, and Empathy/Subjectivity differed by profile, whereas Access Priority remained low and non-significant after correction. Human validation showed moderate endpoint-specific reliability. Conclusions LLM-generated pediatric asthma referral plans varied in financial-access, geographic-access, navigation, SDOH-recognition, and selected tone-related framing. These findings do not establish discriminatory intent, clinical equivalence, downstream harm, or positive synergistic interaction, but support evaluating structural and access-related framing alongside biomedical content in clinical LLM audits.

Read PDF

Similar papers

Open access Aug 2026

Social Status and Clinical Resource Allocation by a Large Language Model: An Evaluation of 30,618 Decisions

Objective: The objective was to quantify whether demographic and social attributes that were irrelevant to stated clinical need, prognosis, and expected benefit altered resource-allocation decisions made by a general-purpose large language model (LLM). Methods: We conducted a cross-sectional audit of the gpt-5-chat-latest API model alias on 8 October 2025, across seven clinical vignettes, generating 30,618 forced-choice comparisons between patient profiles. Profiles varied across a full-factorial combination of eight demographic and social attributes while clinical need, prognosis, and expected benefit were held constant. Forced choices were analyzed using pooled logistic regression with separate Patient A and Patient B attribute terms and vignette-specific position effects; position-averaged odds ratios and position-balanced absolute probabilities were derived from this model. Priority-score differences were analyzed using an analogous linear model. Results: The model showed large position-averaged associations between non-clinical patient attributes and allocation decisions. Indigenous and Black race were associated with substantially higher odds of selection relative to White race (Indigenous: OR 16.48, 95% CI 14.85–18.28; Black: OR 8.07, 95% CI 7.32–8.90), corresponding to position-balanced absolute increases in selection probability of 30.7 and 16.3 percentage points, respectively. Conversely, high-status occupation (OR 0.064, 95% CI 0.058–0.071), friendship with institutional leadership (OR 0.121, 95% CI 0.111–0.131), and major donor status (OR 0.092, 95% CI 0.084–0.101) were associated with markedly lower odds of selection. Choice-score concordance was 95.1%. Conclusions: In this controlled audit, the LLM’s allocation decisions varied substantially according to demographic and social characteristics despite identical stated clinical need, prognosis, and expected benefit. Although some patterns could be interpreted differently under competing ethical frameworks, their implicit and unexplained incorporation into resource-allocation decisions raises concerns regarding transparency, accountability, and clinical governance. Clinical use of LLM-based allocation support should therefore require explicit safeguards and systematic auditing for non-clinical influences.

Unknown authors · 0 citations
Review Open access Aug 2026

Benchmarking publicly accessible large language models for English-language patient-facing acute pancreatitis information: a cross-sectional study of quality, transparency, and readability

Patients increasingly rely on large language models (LLMs) for health information, yet their suitability for decision-critical conditions such as acute pancreatitis remains unclear. Given that acute pancreatitis requires timely symptom recognition, severity assessment, treatment decision-making, recurrence prevention, and follow-up management, LLM-generated information should demonstrate reliability, transparency, and readability. To evaluate the informational quality, visible transparency-related features, and readability of English-language responses generated by five publicly accessible LLMs to standardized patient-facing questions on acute pancreatitis. This cross-sectional benchmark study developed 24 English-language, single-intent questions on acute pancreatitis across six clinical domains using public search intents and guideline-derived decision-critical content. The analysis focused exclusively on English-language patient-facing responses. Each question was submitted once to GPT-5.4 Thinking, DeepSeek-V3.2, Gemini 3.1 Pro, Grok 4.3, and Qwen3.6-Max-Preview, generating 120 responses. Anonymized responses were independently evaluated by two blinded gastroenterologists using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), and Journal of the American Medical Association (JAMA) benchmark criteria. Readability was assessed using six established formulas. A structured response-level safety analysis was added to evaluate factual inaccuracies, clinical hallucinations, and clinical safety-risk severity. Between-model differences were analyzed with Friedman tests followed by post hoc paired Wilcoxon signed-rank tests with Holm correction. Significant between-model differences were observed across all quality, transparency-related, and readability outcomes. Grok 4.3 achieved the highest mean DISCERN, GQS, EQIP, and JAMA scores, reflecting the strongest informational quality profile according to the predefined quality instruments, although visible transparency cues remained limited across all models. In the added safety analysis, factual inaccuracies and clinical hallucinations were each identified in 16 of 120 responses (13.3%), whereas clinical safety-risk signals were identified in 7 responses (5.8%), all of which were adjudicated as score 1 (low risk) under the predefined clinician-rated rubric; no moderate- or high-risk clinical safety event was adjudicated. DeepSeek-V3.2 demonstrated the most favorable readability profile, with the highest Flesch Reading Ease score (44.42 ± 12.27), which nevertheless remained substantially below the recommended threshold of ≥ 80. None of the 120 responses satisfied all six predefined readability targets. All model-level readability distributions differed significantly from recommended thresholds in the direction of poorer readability. Publicly accessible LLMs generated English-language responses on acute pancreatitis with variable informational quality, limited visible transparency cues, and consistently inadequate readability under standardized default public-interface conditions. Because the primary analysis used a single-generation cross-sectional design, model rankings should be interpreted as performance snapshots rather than definitive or temporally stable hierarchies. Higher quality scores did not necessarily establish factual accuracy, clinical safety, patient-education-level readability, or stronger response-level transparency. Current public-interface LLM outputs may serve as clinician-reviewed drafts for patient education but should not function as standalone patient resources. Future AI-based health information systems should strengthen clinical completeness, plain-language communication, visible evidence support, actionability, and clinician-supervised safeguards.

Biao Jiang, Hongxin Sun, Linlin Chen · 0 citations
Review Open access Aug 2026

AI safety evaluation in an underrepresented population: real-world performance of clinical decision support and frontier language models on Medicaid patient messaging triage

No evaluated tool or combination was sufficiently accurate to enable physician-unassisted triage in this setting of patient-initiated text messages in a multi-state Medicaid population.

Sanjay Basu, Sadiq Y. Patel, Parth Sheth et al. · 0 citations
Open access Aug 2026

Quality, readability, and clinical-risk signals of default public-interface LLM responses to vascular and perioperative patient questions: a cross-sectional snapshot

Background Patients increasingly use public large language model chatbot interfaces to seek health information. In vascular disease and perioperative management, default first responses may influence how patients interpret urgent symptoms, antithrombotic medications, procedural choices, and anesthesia-related safety issues. Methods This cross-sectional benchmark study evaluated 110 default first responses returned by five publicly accessible LLM chatbot interfaces during a defined access window on May 28–29, 2026, Beijing time (UTC + 8). Interface names were recorded solely as the public-interface display labels visible at the time of access and should not be interpreted as independently verified API-level model identifiers. Each model was queried with 22 guideline-derived English patient-facing questions, yielding 110 responses. Responses were assessed using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), Journal of the American Medical Association (JAMA) benchmark criteria, used only as a visible metadata/transparency proxy, six readability formulas, an investigator-developed Guideline Concordance Score, and an investigator-developed Potential Clinical-Risk Severity Flag. Differences across the evaluated public-interface response sets were tested using Friedman tests with Holm-adjusted post hoc comparisons. Results Interrater agreement was high for established instruments: DISCERN ICC(A,1) = 0.940, EQIP ICC(A,1) = 0.830, GQS weighted κ = 0.829, and JAMA weighted κ = 0.898. Observed public-interface response performance differed significantly for DISCERN, EQIP, and JAMA criteria (all p < 0.001), and for GQS (p = 0.011). Within this defined late-May 2026 response set, responses returned by the interface displaying the label ‘Grok-4.3’ had the highest observed sample means for DISCERN (58.68), EQIP (81.14), GQS (4.27), and the JAMA visible metadata/transparency proxy score (1.00). No evaluated public-interface response set achieved recommended sixth-grade readability, and no individual response met all six readability thresholds. Responses returned by the interface displaying the label ‘DeepSeek-v4’ had the most favorable observed readability profile within this sampled response set, although the mean FKGL remained 11.53. These observations should not be interpreted as durable model rankings. Conclusion In this English-language benchmark of default first responses from five public LLM interfaces accessed through specific logged-in accounts from a Hong Kong, China IP address during a single late-May 2026 window, 107 of 110 responses received the maximum Guideline Concordance Score, and no response met the prespecified criteria for a high Potential Clinical-Risk Severity Flag. Pronounced ceiling and floor effects preclude conclusions about clinical sufficiency or safety. These findings should not be generalized to non-English use, different health-literacy levels, country-specific emergency-care pathways, other regions or account configurations, or later interface states.

Wei Zhong, Yuanyuan Zhang, Yu Huang et al. · 0 citations
Aug 2026

Clinical safety of large language model responses to matched patient-language and clinician-language Turkish obstetric and gynecologic triage prompts: a model-blinded paired-scenario study.

OBJECTIVE To determine whether presenting matched obstetric and gynecologic triage scenarios as patient-language prompts rather than clinician-language prompts affects expert-rated clinical confidence and the clinical safety of LLM-generated advice. METHODS Thirty obstetric and gynecologic scenarios were presented in matched clinician-language and patient-language Turkish formats to four LLMs. Five specialists independently evaluated 240 responses, generating 1,200 ratings. The primary outcome was the Global Clinical Confidence Score (GCCS; 0-2); five secondary outcomes were rated on 1-5 scales. Associations were examined using ordinal logistic generalized estimating equations adjusted for model and evaluator. RESULTS Clinically reliable responses (GCCS = 2) accounted for 89.3% of clinician-language and 91.3% of patient-language ratings. Patient-language phrasing was not significantly associated with overall GCCS (cumulative odds ratio 0.78, 95% confidence interval 0.57-1.06; p = 0.115), and the language-by-model interaction was not significant (p = 0.422). Patient-language prompts were associated with fewer GCCS = 0 ratings in a binary generalized estimating equations analysis (odds ratio 0.66, 95% confidence interval 0.46-0.96; p = 0.031), although the exact paired McNemar test was not significant (p = 0.096). After false-discovery-rate correction, patient-language prompts had higher evaluator-level triage appropriateness and clinical applicability scores (both adjusted p = 0.028). CONCLUSION No significant difference in overall expert-rated clinical confidence was detected between patient-language and clinician-language prompts.

Onur Ada, Uğurcan Dağlı, E. Bilen et al. · 0 citations
Open access Aug 2026

Same child, different risk: demographic bias in childhood obesity attribution by large language models

Background Large language models (LLMs) are increasingly consulted for pediatric health information, yet their demographic biases remain unsystematically evaluated in pediatric contexts. Objectives To assess bias and variability in childhood obesity risk attribution across seven LLMs (ChatGPT, Claude, DeepSeek, Gemini, GLM, Grok, and Qwen), spanning both Western and Chinese-origin developers; all prompts, including those submitted to the Chinese-origin models, were in English only. Methods A structured prompt-based experimental design was employed across six clinical domains (general obesity risk, dietary pattern, physical activity, sleep, mental health, and genetic predisposition) and six demographic comparison dimensions (sex, three race/ethnicity pairings, socioeconomic status, and urban-rural residence). Seventy-eight unique prompts were submitted to each model in triplicate, yielding 1,638 outputs. Neutral prompts were scored on a five-dimension binary rubric (accuracy, representation, stigmatizing/harmful language, social determinants, cultural fit); comparative prompts were coded for directional risk attribution. Results Claude achieved the highest neutral prompt composite score (mean 3.00 ± 0.91) and GLM the lowest (1.44 ± 0.51); between-model differences were statistically significant (Kruskal–Wallis H = 46.21, p < 0.001). All models achieved a 100% Stigmatizing/Harmful Language pass rate, yet representation and cultural fit were universally weak. Socioeconomic status produced the most consistent attribution pattern (low-income attribution in 40/42 decisions; decision change rate 19.0%). Most models attributed higher obesity risk to Black and Hispanic/Latino children across the majority of domains. Urban–rural attribution showed the greatest cross-model directional inconsistency (decision change rate 52.4%), with Western-origin models favoring rural attribution and Chinese-origin models favoring urban attribution. Conclusions Publicly accessible English-language web-interface outputs from current LLMs showed systematic demographic patterns in pediatric obesity risk attribution, supporting the need for pre-deployment and post-deployment bias auditing before clinical or consumer health use.

Can Wang, Zhendong Liu, Yanyu Jiang et al. · 0 citations