Skip to content

Large Language Models in Oral and Maxillofacial Surgery Triage: A Scoping Review

Aug 2026 · Current Surgery Reports · Vol 14 · 0 citations · 11 references

TL;DR

Large Language Models show potential in their diagnostic accuracy and consequent ability to reduce clinician burden, and may provide the greatest benefit when used to optimise referral quality at source, improving both clinician and potentially LLM triage downstream.

View source

Similar papers

Open access Jul 2026

A comparative evaluation of large language models in diagnosis and treatment planning in restorative dentistry.

LLMs appear to have potential as supplementary information resources for clinicians in restorative dentistry; however, their clinical integration, impact on patient outcomes, and real-world usability remain to be established in future research.

Ebru İrem Teke, A. Borsöken · 0 citations
Review Open access Jul 2026

Patients' perception towards large language models in otorhinolaryngology, head and neck surgery: a single-centre survey

OrL-HNS patients are largely familiar with LLMs and frequently use them, but their trust and confidence regarding health information provided by LLMs alone is limited, suggesting that clinician-led LLM use may be acceptable to many patients.

C. Buhr, A. Blaikie, Harry Smith et al. · 0 citations
Jul 2026

Diagnostic accuracy of large language models in ICOP-based orofacial pain diagnosis: A comparative study.

OBJECTIVE To compare the diagnostic performance of ChatGPT 5.5, Claude Opus 4.1, Gemini 3 Flash, and Grok 4 in International Classification of Orofacial Pain (ICOP)-based clinical scenarios. METHODS Thirty ICOP diagnoses were randomly selected, and corresponding clinical scenarios were manually developed. Each scenario was submitted to all models using standardized prompts in independent sessions. Two blinded evaluators assessed primary diagnosis accuracy, subclassification accuracy, clinical interpretation, and management recommendations. RESULTS  Overall performance differed significantly among models (p < .001). Grok 4 achieved the highest total score and outperformed the other models. No significant differences were found among ChatGPT 5.5, Gemini 3 Flash, and Claude Opus 4.1. Subclassification accuracy was consistently lower than primary diagnosis accuracy, while management recommendations did not differ significantly. CONCLUSION LLM performance varied across ICOP-based scenarios. Although Grok 4 showed the highest diagnostic concordance, current LLMs should support, not replace, clinician judgment.

M. S. Şimşek, Enis Esen, M. Koparal · 0 citations
Review Open access Aug 2026

Evaluation of a large language model for clinician-facing preoperative cost-communication preparation in total knee arthroplasty

Background and aims Costs associated with total knee arthroplasty (TKA) may affect treatment preparation, expectation management, and postoperative care planning. Previous large language model (LLM) studies have focused mainly on medical question answering, patient education, and clinical decision support, whereas their performance in clinician-facing preoperative cost-communication preparation remains unclear. This study evaluated an LLM using real-world clinical records from two hospitals in an expert-referenced offline evaluation. Methods Preoperative medical records of 80 patients who underwent primary unilateral TKA at two hospitals in China from January to May 2026 were included. A structured expert-panel process was used to develop a preoperative cost-communication framework comprising 4 dimensions and 17 clinical cues and to establish case-level minimum necessary communication items (CL-MNCIs) for each case. Task 1 assessed identification of the 17 cues. Task 2 assessed CL-MNCI coverage and classified all generated items according to case relevance, medical-record support, redundancy, and safety. Twenty-four cases were non-randomly selected by CL-MNCI count for three repeated-generation runs. Results Task 1 comprised 1,360 case–label classification units. Micro-precision, micro-recall, micro-F1, accuracy, MCC, and macro-F1 were 0.901, 0.888, 0.894, 0.948, 0.860, and 0.840, respectively. Experts established 569 CL-MNCIs, of which 483 were covered, yielding an overall coverage rate of 84.9%; complete coverage was achieved in 17 cases. The LLM generated 716 items, including 12 safety events across 9 cases, 483 items matching CL-MNCIs, 132 record-supported supplementary items, 50 redundant items, and 39 items with insufficient record support. Overall, 665 items (92.9%) were case-relevant and record-supported, although this proportion included redundant content. Pairwise Jaccard similarity for covered CL-MNCI sets ranged from 0.813 to 0.823. Of 180 CL-MNCIs, 126 (70.0%) were covered in all three runs, and case-level agreement in safety classification ranged from 87.5 to 95.8%. Conclusion The LLM showed offline potential for identifying cost-communication cues and generating clinician-facing preparation checklists for TKA, but content omissions, quality variation, safety risks, and substantive cross-run variability remained. Its use should be limited to clinician-reviewed communication preparation and should not replace direct patient cost disclosure or professional judgment.

Zebing Ma, Liping Xue, Gonghui Jian et al. · 0 citations
Open access Mar 2026

Can Large Language Models Identify When an Upper Extremity Problem Needs Nonurgent Attention? An Assessment of Multiple LLM Chatbots.

IntroductionSeeking emergency care regarding musculoskeletal sensations is far more prevalent than limb or life-threatening pathophysiology. We studied the ability of an LLM to distinguish between urgent and nonurgent upper extremity symptoms and provide an accurate diagnosis.MethodsFive LLMs (ChatGPT-4, ChatGPT-4o, Co-Pilot, Gemini, and PerplexityChat) were presented with descriptions of seven urgent and seven nonurgent symptom scenarios written below a sixth grade reading level. LLM responses were identified as appropriate if immediate urgent medical attention was recommended after an initial and ongoing inquiry ("What additional information do you need to diagnose my condition?"). Diagnoses provided were identified as correct, partially correct, or incorrect. The analysis was repeated 24 months later with the current LLM versions and results were compared.ResultsLLMs discerned nonurgent conditions with 97% positive predictive value (PPV) and an 89% negative predictive value (NPV) on initial query, which improved to 96% and 97% respectively after ongoing inquiry. Compartment syndrome was misidentified as nonurgent in 80% of scenarios on initial inquiry, although four of five LLMs corrected on continued inquiry. Diagnosis was correct or partially correct for 115 of 150 (82%) on initial inquiry. An updated analysis 24 months later demonstrated marked improvement in LLM ability to identify emergencies with 100% PPV and 95% NPV on initial query and 98% NPV after ongoing inquiry.ConclusionThe finding that LLMs can distinguish urgent from nonurgent upper extremity conditions suggests that artificial intelligence tools could help reduce unnecessary use of high-cost emergency services, allowing those resources to be reserved for patients who require timely care.

Jefferson Hunter, David Ring, Prakash Jayakumar · 0 citations

Related blog posts