Jul 2026· Frontiers in Digital Health· Vol 8· 0 citations· 26 references
Medicine
TL;DR
OrL-HNS patients are largely familiar with LLMs and frequently use them, but their trust and confidence regarding health information provided by LLMs alone is limited, suggesting that clinician-led LLM use may be acceptable to many patients.
Abstract
Objectives Large language models (LLMs) are increasingly discussed for use in clinical practice. Beyond their performance, patients' acceptance is crucial for their implementation. We investigated ORL-HNS patients' familiarity with AI/LLMs, use patterns, and trust in LLM-based medical information and recommendations. Methods In this single-centre prospective survey at a German university hospital, ORL-HNS patients with and without malignant disease completed a 15-item questionnaire. Results A total of 123 patients, 20 (16%) with and 103 (84%) without malignant disease, participated in the study. Most patients were familiar with the term AI (96%, n = 118) and LLMs (78%, n = 96). Overall, 72/123 (59%) reported using LLMs. One third (33%, n = 40) retrieved “Health information”, rating the LLMs with median Likert scores for comprehensibility 5 [IQR 4, 6], conciseness 5 [IQR 3, 6] and coherence 5 [IQR 3, 6]. However, perceived medical accuracy received a median rating of 4 [IQR 3, 5], significantly lower than comprehensibility (p < 0.05). With respect to the confidence in the recommendations exclusively by LLMs [median 2 (IQR 2, 3.5)] received significantly lower ratings than doctors [median 5 (IQR 5, 6)] and doctors also using LLMs [median 5 (IQR 4, 6)], p < 0.0001 respectively. Conclusion ORL-HNS patients are largely familiar with LLMs and frequently use them, but their trust and confidence regarding health information provided by LLMs alone is limited. Patients show the greatest confidence in doctors' recommendations. Yet they reported similar confidence in physician recommendations and physician recommendations supported by LLMs, suggesting that clinician-led LLM use may be acceptable to many patients.
LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.
Hetal Lad, Emily S Kwon, Ayushi Chadha et al.· Journal of Otorhinolaryngolo...· 0 citations
Large Language Models show potential in their diagnostic accuracy and consequent ability to reduce clinician burden, and may provide the greatest benefit when used to optimise referral quality at source, improving both clinician and potentially LLM triage downstream.
K. Surendran, I. Aziz, Glyndwr Jenkins· Current Surgery Reports· 0 citations
OBJECTIVES
Systematic literature reviews (SLRs) are time-intensive and resource-consuming. While large language models (LLMs) have shown promise in encoding clinical knowledge, evidence for their performance on complex text analysis, necessary for SLRs in otolaryngology, remains limited. In this proof-of-concept study, we investigate an LLM's performance in screening relevant articles for a SLR with the novel screening of title and abstracts, reevaluation, and full-text review (STARR) protocol.
METHODS
ENTGPT (based on GPT-4o) was compared to two human reviewers in article inclusion/exclusion decisions using the traditional and STARR screening protocols. ENTGPT was provided with inclusion and exclusion criteria, titles, abstracts, and full texts (if available) of the 850 articles retrieved in the original search. The model's decisions were compared to those made by two human reviewers.
RESULTS
ENTGPT, using the STARR protocol, achieved 99.87% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 95% sensitivity. When using the traditional protocol, sensitivity declined to 35%. ENTGPT, using the traditional protocol, achieved 99.47% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 35% sensitivity. When using the STARR protocol, accuracy improved to 99.87% and sensitivity markedly increased to 95%.
CONCLUSIONS
ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols. This performance suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers.
LEVEL OF EVIDENCE
N/A.
Akash Kapoor, Ben Baranker, I. Alter et al.· The Laryngoscope· 0 citations
This study aimed to compare the educational performance of six mainstream LLMs for neuromyelitis optica spectrum disorder (NMOSD) and evaluated patient satisfaction during real-world interactions.
This study was conducted from March to April 2026. In the first Phase, Twenty NMOSD-related questions derived from clinical guidelines and patient concerns were submitted to six LLMs (ChatGPT-5.4, Gemini-3.1-pro, Claude-4.6-Sonnet, DeepSeek-3.2, Kimi-2.5, and Qwen-3.5-plus). Responses were anonymized and independently evaluated by three neuro-ophthalmology specialists using Likert framework assessing accuracy, completeness, readability, safety, and humanity. Inter-rater reliability was assessed using the intraclass correlation coefficient (ICC). In the second Phase, the three best-performing models were subsequently evaluated through real-world interactions with ten NMOSD patients, and satisfaction scores were analyzed using linear mixed-effects models.
A total of 120 chatbot responses were evaluated. With a comprehensive evaluation, significant differences were observed across all assessment domains. Gemini-3.1-pro achieved the highest scores for accuracy and safety, while Qwen-3.5-plus demonstrated superior completeness and humanity. DeepSeek-3.2 generated the most accessible responses, exhibiting the lowest reading difficulty score. However, its completeness advantage should be interpreted with caution, as it may be partially influenced by its longer response length. Inter-rater reliability was good, with single-measure ICC values ranging from 0.535 to 0.759, and average-measure ICC values ranging from 0.775 to 0.904. In patient interactions, Qwen-3.5-plus achieved the highest satisfaction score, significantly outperforming Gemini-3.1-pro and DeepSeek-3.2. Although all LLMs demonstrate superior performance in patient education, the real-world interaction with patient needs to pay attention.
LLMs demonstrate considerable potential for NMOSD patient education but exhibit variability across educational dimensions, and require further validation in larger cohorts. These findings highlight the importance of selecting LLMs according to specific patient education goals and underscore the importance of clinicians in rare disease counseling.
Chen Li, Yuting Hu, Xiaoyan Wang et al.· Frontiers in Physiology· 0 citations
Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.
Javier De la Torre-Trillo, Albert Munuera, M. D. Ureña et al.· Clinical and Translational O...· 0 citations