Aug 2026· Journal of Otorhinolaryngology Hearing and Balance Medicine· 0 citations· 24 references
TL;DR
LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.
Abstract
Background/Objectives: As patients increasingly rely on large language models (LLMs) for Chronic Rhinosinusitis (CRS) diagnosis, surgical candidacy, and perioperative care, evaluating the accuracy of LLM-generated information against established clinical practice guidelines for surgical management of CRS is essential. Methods: ChatGPT, Google AI, Google Gemini, and Grok were queried using a 21-question guideline-mapped prompt set (long) and a single patient-focused prompt (short). Two physician reviewers independently scored responses using a 3-point rubric across 21 fields. Primary outcomes were guideline-concordant scores; secondary outcomes included readability measured with the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater reliability (IRR) was assessed using the intraclass correlation coefficient (ICC). Analyses were performed in SPSSv31. Results: Guideline concordance ranged from 55.36% to 77.98% (p > 0.05), highest for Grok (77.98%, 95% CI 63.67–92.28), followed by Google Gemini (66.67%, 95% CI 35.42–97.91), ChatGPT (55.95%, 95% CI 29.25–82.65), and Google AI (55.36%, 95% CI 29.97–80.75), with Grok significantly outperforming both ChatGPT and Google AI. Prompt structure significantly affected scores. Long-form prompting resulted in higher guideline concordance scores than short-form prompting (+26.19, p < 0.001). The CPG demonstrated a more readable structure, with a higher FRE (44.1), exceeding scores generated by Grok (31.7), ChatGPT (39.9), Gemini (39.2), and Google AI (32.5). In contrast, the CPG was a higher reading grade level (FKGL score of 11.7) than Grok (11.4), Gemini (10.5), and ChatGPT (10.3), but was lower in reading grade compared to Google AI, which produced the highest FKGL score (12.5). IRR was high (ICC = 0.961). Conclusions: LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms. While LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability. However, longer prompt structure meaningfully influenced output quality, highlighting how a user’s ability to frame precise prompts is critical to obtaining accurate information.
LLM-generated responses to parent-oriented reformulations of Pediatric Infectious Diseases Society/Infectious Diseases Society of America guideline recommendations showed high guideline concordance, but selected item-level discordance persisted.
Ahmet Murat Çörekci, Belen Ateş, Orkun Dinç et al.· The Pediatric Infectious Dis...· 0 citations
This study aimed to compare the accuracy, completeness, and readability of responses generated by three large language models (LLMs)—ChatGPT-4.0 (OpenAI), Microsoft CoPilot, and Google Gemini—regarding the treatment and management of idiopathic congenital talipes equinovarus (ICTEV) using the Ponseti method.
Fifteen frequently asked questions were selected from pediatric orthopedic center websites, Google Trends analysis, and clinical experience. Each question was submitted verbatim in a new session to the three LLMs within 24 h. Seven board-certified pediatric orthopedic surgeons, blinded to the source, rated responses for accuracy (5-point Likert scale) and completeness (3-point Likert scale). Readability was assessed using the Flesch–Kincaid grade level. Mean scores ± standard deviation were calculated, and inter-rater reliability was estimated using the intraclass correlation coefficient (ICC). Group differences were tested with ANOVA and chi-squared tests (
p
< 0.05).
A total of 45 responses were evaluated. Gemini achieved the highest mean accuracy (4.1 ± 0.8), followed by CoPilot (3.6 ± 0.8) and ChatGPT-4.0 (3.3 ± 0.9), with significant differences among models (
p
< 0.001). Completeness ratings also differed significantly (
p
< 0.001). Readability analysis showed that ChatGPT produced shorter, more readable text, while Gemini generated longer, more complex responses. Inter-rater reliability was substantial for accuracy (ICC 0.715) and completeness (ICC 0.710).
Google Gemini outperformed ChatGPT-4.0 and CoPilot in accuracy and comprehensiveness for ICTEV management using the Ponseti method. However, its complex responses may limit accessibility, whereas ChatGPT-4.0 offered more readable but less detailed answers.
IV
A. Vescio, G. Testa, M. Sapienza et al.· Journal of Children's Orthop...· 0 citations
Purpose To evaluate the guideline knowledge alignment of two large language models (LLMs), GPT‐5.5 Instant and DeepSeek‐V4, and to determine their preclinical reliability as reference tools in refractive surgery. Methods Using the 38 evidence‐based recommendations of the international keratorefractive lenticule extraction (KLEx) guidelines as the gold standard, both LLMs were evaluated in their default configurations. Two ophthalmologists independently assessed the clinical safety and medical accuracy of the model responses using a 5‐point Likert scale (1–5 points). Agreement between each model’s recommendation strength and the guideline was quantified by intraclass correlation coefficient (ICC), structural reliability by the DISCERN scale, and readability by the Flesch Reading Ease (FRE) and Flesch–Kincaid Grade Level (FKGL) indices. Results Likert ratings did not differ between GPT‐5.5 Instant (4.96 ± 0.206) and DeepSeek‐V4 (4.89 ± 0.385; p = 0.134). Across 114 independent generations, the ICC for agreement with the guideline was 0.884 (95% CI, 0.836–0.918) for GPT‐5.5 Instant and 0.739 (0.640–0.813) for DeepSeek‐V4 (both p < 0.001). DISCERN scores were 70.18 ± 5.16 and 68.05 ± 5.41 (p = 0.085); FRE, 9.93 ± 8.33 and 3.74 ± 5.24 (p < 0.001); and FKGL, 16.07 ± 2.21 and 18.95 ± 1.96 (p < 0.001). Conclusion Both models aligned closely with the guidelines, with significant concordance in GRADE‐based recommendation strength. They may serve as preclinical reference tools in refractive surgery but not as a substitute for specialist judgment.
Hong-Xia Lu, Yan Huo, Rui-Si Xie et al.· Journal of Ophthalmology· 0 citations
This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management.
Large language models (LLMs) are increasingly used for patient health education, yet the readability and educational quality of LLM-generated information on trigeminal neuralgia (TN) have been insufficiently evaluated. This cross-sectional benchmarking study compared TN educational content generated by five publicly available LLMs (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5). Twenty frequently asked TN questions covering basic disease knowledge, etiology/risk factors, diagnosis, treatment, and prevention/rehabilitation were presented to each model using standardized prompts. Readability was assessed using seven established indices, educational suitability using the Patient Education Materials Assessment Tool for Understandability and Actionability (PEMAT), and overall information quality using the Global Quality Score (GQS). Two clinical experts independently evaluated all responses, with disagreements resolved by a senior adjudicator. Statistical analyses compared model performance, thematic differences, and correlations among the evaluation metrics. Significant differences were observed among the models for readability, PEMAT, and GQS scores. GPT-5 generated the most linguistically complex responses but achieved the highest ratings for educational suitability and information quality. In contrast, Wenxin Yiyan produced the most readable text but generally scored lower on PEMAT and GQS. Content category influenced readability, with prevention/rehabilitation and etiology/risk-factor topics being more difficult to read, whereas PEMAT and GQS remained relatively consistent across themes. Readability indices showed strong internal consistency and weak-to-moderate positive correlations with PEMAT and GQS, while PEMAT and GQS demonstrated a moderate positive correlation. These findings suggest that model selection influences expert-rated educational suitability and overall information quality, whereas topic complexity primarily affects readability. Because patient comprehension, satisfaction, trust, health outcomes, factual accuracy, and clinical safety were not evaluated, these results should be interpreted as an expert-rated benchmarking analysis rather than evidence of clinical readiness.
Hao Wei, Sisi Sun, Mingxin Liu et al.· Journal of Visualized Experi...· 0 citations