Skip to content
Review

Guideline Concordance of Large Language Model Responses to Parent-oriented Guideline Prompts About Pediatric Acute Bacterial Arthritis.

Jul 2026 · The Pediatric Infectious Disease Journal · 0 citations · 32 references
Medicine

TL;DR

LLM-generated responses to parent-oriented reformulations of Pediatric Infectious Diseases Society/Infectious Diseases Society of America guideline recommendations showed high guideline concordance, but selected item-level discordance persisted.

Abstract

Background

Large language models (LLMs) are increasingly used by patients and caregivers to obtain medical information. In pediatric acute bacterial arthritis, non-guideline-concordant information may be clinically important because timely diagnosis and management are essential. This study evaluated the concordance of LLM-generated responses to parent-oriented reformulations of Pediatric Infectious Diseases Society/Infectious Diseases Society of America (PIDS/IDSA) guideline recommendations.

Methods

In this exploratory cross-sectional benchmarking study, 27 PIDS/IDSA guideline-derived recommendations and good practice statements were reformulated into standardized parent-oriented prompts. The same prompts were submitted to GPT-5.4 Thinking, Gemini 3 Thinking and Claude 4.6 Sonnet through browser-based interfaces on April 12, 2026. Responses were anonymized and independently assessed by 3 blinded reviewers. Each response was classified as concordant or discordant; final classifications were determined by majority decision. Interrater agreement was assessed using Fleiss' kappa, and model differences were evaluated using Cochran's Q test.

Results

Overall, 75 of 81 responses (92.6%) were concordant with PIDS/IDSA recommendations. Gemini 3 Thinking achieved concordance in 27/27 responses (100.0%), Claude 4.6 Sonnet in 25/27 (92.6%) and GPT-5.4 Thinking in 23/27 (85.2%). Cochran's Q test showed no significant difference among models (Q = 4.800, df = 2, P = 0.091). No unsupported or hallucinated content was identified. Interrater agreement was moderate (κ = 0.580; 95% CI, 0.454-0.706; P < 0.001).

Conclusions

LLMs showed high guideline concordance, but selected item-level discordance persisted. Because these models may have been trained on guideline-derived content, high concordance should not be equated with independent clinical reasoning. These tools may support caregiver-oriented education but should not replace clinician assessment or guideline-based care.

View source

Similar papers

Review Open access Aug 2026

Benchmarking Large Language Model Responses Against Surgical Clinical Practice Guidelines for Chronic Rhinosinusitis: The Importance of User Prompts

LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.

Hetal Lad, Emily S Kwon, Ayushi Chadha et al. · 0 citations
Aug 2026

Clinical safety of large language model responses to matched patient-language and clinician-language Turkish obstetric and gynecologic triage prompts: a model-blinded paired-scenario study.

OBJECTIVE To determine whether presenting matched obstetric and gynecologic triage scenarios as patient-language prompts rather than clinician-language prompts affects expert-rated clinical confidence and the clinical safety of LLM-generated advice. METHODS Thirty obstetric and gynecologic scenarios were presented in matched clinician-language and patient-language Turkish formats to four LLMs. Five specialists independently evaluated 240 responses, generating 1,200 ratings. The primary outcome was the Global Clinical Confidence Score (GCCS; 0-2); five secondary outcomes were rated on 1-5 scales. Associations were examined using ordinal logistic generalized estimating equations adjusted for model and evaluator. RESULTS Clinically reliable responses (GCCS = 2) accounted for 89.3% of clinician-language and 91.3% of patient-language ratings. Patient-language phrasing was not significantly associated with overall GCCS (cumulative odds ratio 0.78, 95% confidence interval 0.57-1.06; p = 0.115), and the language-by-model interaction was not significant (p = 0.422). Patient-language prompts were associated with fewer GCCS = 0 ratings in a binary generalized estimating equations analysis (odds ratio 0.66, 95% confidence interval 0.46-0.96; p = 0.031), although the exact paired McNemar test was not significant (p = 0.096). After false-discovery-rate correction, patient-language prompts had higher evaluator-level triage appropriateness and clinical applicability scores (both adjusted p = 0.028). CONCLUSION No significant difference in overall expert-rated clinical confidence was detected between patient-language and clinician-language prompts.

Onur Ada, Uğurcan Dağlı, E. Bilen et al. · 0 citations
Aug 2026

Large Language Models as Clinical Support Tools in Drug Information Services: Performance Comparison With Pharmacists.

LLMs show potential as tools for preliminary drug information retrieval and rapid responses generation in drug information services, however, variable concordance and persistent limitations in citation credibility indicate the need for continued pharmacist oversight.

Nuntapong Boonrit, Najwa Bin-Useng, Aphichaya Sirijariyawat et al. · 0 citations
Jul 2026

The illusion of competence: Evaluating the clinical reasoning of large language models in pediatric gastroenterology.

While current LLMs demonstrate good diagnostic pattern recognition in PGHN, reproducible and potentially life-threatening failures in pharmacological reasoning and reference accuracy create a dangerous illusion of competence.

Y. Ergen, S. Teke, E. G. Başaran et al. · 0 citations
Open access Aug 2026

Mapping Gaps and Improvement Targets in Large Language Model-Generated Melanoma Patient Education in a Non-English Setting

Objective: Large language models (LLMs) are increasingly being used to develop medical education materials; however, it remains unclear how reliable, readable, or guideline-compliant the content generated by these models is for non-English-speaking patient groups. We evaluated the quality of Turkish melanoma patient education texts generated by seven frontier LLMs. Methods: A standardized 22-item Turkish prompt, built from international melanoma guidelines, was put to seven models in zero-shot sessions: ChatGPT 4.0 Turbo, Gemini 2.0 Flash, Claude 3.7 Sonnet, Grok 3, Qwen 2.5 Plus, DeepSeek R1, and Mistral Large 2. Each output was rated for readability (Ateşman Index), understandability, how clearly medical terminology was explained, scientific reliability (DISCERN instrument), empathy, and adherence to a 31-item guideline-based checklist. Model comparisons were summarized descriptively, using model-level absolute scores, score ranges, and rankings. Results: Model performance differed across readability, understandability, reliability, empathy, and guideline-adherence domains. DeepSeek R1 led on both readability (81.6) and understandability (23.5/25). Guideline adherence was strongest for Grok 3 and DeepSeek R1, at 96.8% and 93.5%, respectively, and Grok 3, DeepSeek R1, and Gemini 2.0 Flash each scored above 90% on the normalized total DISCERN measure. DeepSeek R1 also recorded the highest empathy score (90%). Gemini 2.0 Flash had the lowest readability score and produced the longest output (Ateşman 65.8). None of the models provided citations or verifiable sources, so every model received the lowest possible DISCERN Source Reliability score; Mistral Large 2 showed the weakest overall performance. Conclusion: How well large language models (LLM) handle Turkish melanoma patient education varies widely from one model to the next. A few produced text that was clear, empathetic, and reasonably guideline-concordant, but the lack of verifiable citations and uneven guideline coverage remain genuine limitations. These findings suggest that LLM-generated Turkish melanoma materials may be useful as preliminary educational drafts.

Niyazi Çetin, A. Atılan · 0 citations
Open access Aug 2026

Evaluating search-enabled large language model interfaces for mpox public health consultation: a guideline-based comparative study

The evaluated search-enabled LLM interfaces showed heterogeneous performance across safety, accuracy, empathy, reliability/information quality, and readability, which support the need for guideline-based evaluation, source transparency, readability optimization, and robust safety safeguards when such interfaces are evaluated or considered for mpox-related public health consultation.

Qi-Qi Zheng, Ru Chen, Ming-Ming Cai et al. · 0 citations