Clinical safety of large language model responses to matched patient-language and clinician-language Turkish obstetric and gynecologic triage prompts: a model-blinded paired-scenario study.
Aug 2026· European Journal of Obstetrics, Gynecology, and Reproductive Biology· Vol 326, pp.
115351
· 0 citations· 19 references
Medicine
Abstract
Objective
To determine whether presenting matched obstetric and gynecologic triage scenarios as patient-language prompts rather than clinician-language prompts affects expert-rated clinical confidence and the clinical safety of LLM-generated advice.
Methods
Thirty obstetric and gynecologic scenarios were presented in matched clinician-language and patient-language Turkish formats to four LLMs. Five specialists independently evaluated 240 responses, generating 1,200 ratings. The primary outcome was the Global Clinical Confidence Score (GCCS; 0-2); five secondary outcomes were rated on 1-5 scales. Associations were examined using ordinal logistic generalized estimating equations adjusted for model and evaluator.
Results
Clinically reliable responses (GCCS = 2) accounted for 89.3% of clinician-language and 91.3% of patient-language ratings. Patient-language phrasing was not significantly associated with overall GCCS (cumulative odds ratio 0.78, 95% confidence interval 0.57-1.06; p = 0.115), and the language-by-model interaction was not significant (p = 0.422). Patient-language prompts were associated with fewer GCCS = 0 ratings in a binary generalized estimating equations analysis (odds ratio 0.66, 95% confidence interval 0.46-0.96; p = 0.031), although the exact paired McNemar test was not significant (p = 0.096). After false-discovery-rate correction, patient-language prompts had higher evaluator-level triage appropriateness and clinical applicability scores (both adjusted p = 0.028).
Conclusion
No significant difference in overall expert-rated clinical confidence was detected between patient-language and clinician-language prompts.
LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.
Hetal Lad, Emily S Kwon, Ayushi Chadha et al.· Journal of Otorhinolaryngolo...· 0 citations
While current LLMs demonstrate good diagnostic pattern recognition in PGHN, reproducible and potentially life-threatening failures in pharmacological reasoning and reference accuracy create a dangerous illusion of competence.
Y. Ergen, S. Teke, E. G. Başaran et al.· Journal of Pediatric Gastroe...· 0 citations
Current-generation LLMs demonstrated superior performance in diagnostic accuracy and management score compared with healthcare professionals and medical students on standardized pediatric infectious disease cases, and these findings support their potential role as clinical decision support tools.
Sait Ramazan Gülbay, Muhammed Yusuf Ozan Avcı, Muhammed Nezih Koç et al.· Journal of Pediatric Infecti...· 0 citations
LLM-generated responses to parent-oriented reformulations of Pediatric Infectious Diseases Society/Infectious Diseases Society of America guideline recommendations showed high guideline concordance, but selected item-level discordance persisted.
Ahmet Murat Çörekci, Belen Ateş, Orkun Dinç et al.· The Pediatric Infectious Dis...· 0 citations
This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management.
Large language models should be regarded as adjunctive, reviewable decision-support tools rather than replacements for MDTs, and GPT-5.4 Thinking showed some concordance with MDT decisions, but strict two-stage overall concordance remained limited.
Yuanze Wei, Yulong Tian, Xiaodong Liu et al.· European Journal of Surgical...· 0 citations