Skip to content
Review Open access

Patients' perception towards large language models in otorhinolaryngology, head and neck surgery: a single-centre survey

Jul 2026 · Frontiers in Digital Health · Vol 8 · 0 citations · 26 references
Medicine

TL;DR

OrL-HNS patients are largely familiar with LLMs and frequently use them, but their trust and confidence regarding health information provided by LLMs alone is limited, suggesting that clinician-led LLM use may be acceptable to many patients.

Abstract

Objectives Large language models (LLMs) are increasingly discussed for use in clinical practice. Beyond their performance, patients' acceptance is crucial for their implementation. We investigated ORL-HNS patients' familiarity with AI/LLMs, use patterns, and trust in LLM-based medical information and recommendations. Methods In this single-centre prospective survey at a German university hospital, ORL-HNS patients with and without malignant disease completed a 15-item questionnaire. Results A total of 123 patients, 20 (16%) with and 103 (84%) without malignant disease, participated in the study. Most patients were familiar with the term AI (96%, n = 118) and LLMs (78%, n = 96). Overall, 72/123 (59%) reported using LLMs. One third (33%, n = 40) retrieved “Health information”, rating the LLMs with median Likert scores for comprehensibility 5 [IQR 4, 6], conciseness 5 [IQR 3, 6] and coherence 5 [IQR 3, 6]. However, perceived medical accuracy received a median rating of 4 [IQR 3, 5], significantly lower than comprehensibility (p < 0.05). With respect to the confidence in the recommendations exclusively by LLMs [median 2 (IQR 2, 3.5)] received significantly lower ratings than doctors [median 5 (IQR 5, 6)] and doctors also using LLMs [median 5 (IQR 4, 6)], p < 0.0001 respectively. Conclusion ORL-HNS patients are largely familiar with LLMs and frequently use them, but their trust and confidence regarding health information provided by LLMs alone is limited. Patients show the greatest confidence in doctors' recommendations. Yet they reported similar confidence in physician recommendations and physician recommendations supported by LLMs, suggesting that clinician-led LLM use may be acceptable to many patients.

Read PDF

Similar papers

Review Open access Aug 2026

Benchmarking Large Language Model Responses Against Surgical Clinical Practice Guidelines for Chronic Rhinosinusitis: The Importance of User Prompts

LLMs demonstrated similar guideline concordance, suggesting patients can expect comparable accuracy across platforms, and LLMs generally improved FKGL scores compared to the AAO-HNS CPG, they demonstrated lower FRE scores, indicating mixed results on overall readability.

Hetal Lad, Emily S Kwon, Ayushi Chadha et al. · 0 citations
#small language model Review Aug 2026

Large Language Models in Oral and Maxillofacial Surgery Triage: A Scoping Review

Large Language Models show potential in their diagnostic accuracy and consequent ability to reduce clinician burden, and may provide the greatest benefit when used to optimise referral quality at source, improving both clinician and potentially LLM triage downstream.

K. Surendran, I. Aziz, Glyndwr Jenkins · 0 citations
Review Aug 2026

ENTGPT: Applying Large Language Models to Systematic Review Screening With the Novel STARR Protocol.

OBJECTIVES Systematic literature reviews (SLRs) are time-intensive and resource-consuming. While large language models (LLMs) have shown promise in encoding clinical knowledge, evidence for their performance on complex text analysis, necessary for SLRs in otolaryngology, remains limited. In this proof-of-concept study, we investigate an LLM's performance in screening relevant articles for a SLR with the novel screening of title and abstracts, reevaluation, and full-text review (STARR) protocol. METHODS ENTGPT (based on GPT-4o) was compared to two human reviewers in article inclusion/exclusion decisions using the traditional and STARR screening protocols. ENTGPT was provided with inclusion and exclusion criteria, titles, abstracts, and full texts (if available) of the 850 articles retrieved in the original search. The model's decisions were compared to those made by two human reviewers. RESULTS ENTGPT, using the STARR protocol, achieved 99.87% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 95% sensitivity. When using the traditional protocol, sensitivity declined to 35%. ENTGPT, using the traditional protocol, achieved 99.47% accuracy in article classification compared with human reviewers (95% CI: 0.99-1.0), including 100% specificity and 35% sensitivity. When using the STARR protocol, accuracy improved to 99.87% and sensitivity markedly increased to 95%. CONCLUSIONS ENTGPT accurately replicated human reviewers in article selection and data extraction for an otolaryngology SLR using the STARR and traditional protocols. This performance suggests that LLMs could be employed to significantly streamline the SLR process, potentially saving substantial time and resources for researchers. LEVEL OF EVIDENCE N/A.

Akash Kapoor, Ben Baranker, I. Alter et al. · 0 citations
Open access Aug 2026

Patient education for neuromyelitis optica spectrum disorder using large language models: combining expert assessment and real-world patient interaction

This study aimed to compare the educational performance of six mainstream LLMs for neuromyelitis optica spectrum disorder (NMOSD) and evaluated patient satisfaction during real-world interactions. This study was conducted from March to April 2026. In the first Phase, Twenty NMOSD-related questions derived from clinical guidelines and patient concerns were submitted to six LLMs (ChatGPT-5.4, Gemini-3.1-pro, Claude-4.6-Sonnet, DeepSeek-3.2, Kimi-2.5, and Qwen-3.5-plus). Responses were anonymized and independently evaluated by three neuro-ophthalmology specialists using Likert framework assessing accuracy, completeness, readability, safety, and humanity. Inter-rater reliability was assessed using the intraclass correlation coefficient (ICC). In the second Phase, the three best-performing models were subsequently evaluated through real-world interactions with ten NMOSD patients, and satisfaction scores were analyzed using linear mixed-effects models. A total of 120 chatbot responses were evaluated. With a comprehensive evaluation, significant differences were observed across all assessment domains. Gemini-3.1-pro achieved the highest scores for accuracy and safety, while Qwen-3.5-plus demonstrated superior completeness and humanity. DeepSeek-3.2 generated the most accessible responses, exhibiting the lowest reading difficulty score. However, its completeness advantage should be interpreted with caution, as it may be partially influenced by its longer response length. Inter-rater reliability was good, with single-measure ICC values ranging from 0.535 to 0.759, and average-measure ICC values ranging from 0.775 to 0.904. In patient interactions, Qwen-3.5-plus achieved the highest satisfaction score, significantly outperforming Gemini-3.1-pro and DeepSeek-3.2. Although all LLMs demonstrate superior performance in patient education, the real-world interaction with patient needs to pay attention. LLMs demonstrate considerable potential for NMOSD patient education but exhibit variability across educational dimensions, and require further validation in larger cohorts. These findings highlight the importance of selecting LLMs according to specific patient education goals and underscore the importance of clinicians in rare disease counseling.

Chen Li, Yuting Hu, Xiaoyan Wang et al. · 0 citations
Review Aug 2026

ChatGPT-4o as a decision-support tool in a urological tumour board: a prospective evaluation.

Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.

Javier De la Torre-Trillo, Albert Munuera, M. D. Ureña et al. · 0 citations