Skip to content
Open access

Artificial intelligence meets pediatric orthopedics: A comparative analysis of ChatGPT-4o, Gemini 2.0, and Claude 3.5 in detecting supracondylar humeral fractures

Jul 2026 · PLoS ONE · Vol 21, pp. e0353782 - e0353782 · 0 citations · 27 references
Medicine

TL;DR

Current LLMs demonstrate limited capability as independent diagnostic tools for pediatric supracondylar humeral fractures and require specialized pediatric training before clinical implementation, however, their potential as assistive tools for triage and assessment warrants further development of pediatric-specific models.

Abstract

Background Supracondylar humeral fractures constitute 10–16% of pediatric skeletal injuries, requiring timely diagnosis to prevent neurovascular complications. Developmental variations in pediatric bone structures pose diagnostic challenges for clinicians. This study evaluated three next-generation large language models (LLMs) (ChatGPT-4o, Gemini 2.0, Claude 3.5) for detecting pediatric supracondylar humeral fractures and their classification according to the Gartland system. Methods This retrospective observational study included 300 pediatric patients (150 with supracondylar humeral fractures confirmed by expert consensus, 150 without fractures) aged 2–10 years presenting to the Emergency Department of the Bilkent City Hospital (October 2022-January 2025). Two-view elbow radiographs were presented to each LLM three times on different days. Diagnostic accuracy was evaluated using overall accuracy (all three responses correct), strict accuracy (≥2 correct responses), and ideal accuracy (≥1 correct response). Response consistency was assessed using Fleiss’ Kappa coefficient. Fractures were classified according to modified Gartland criteria. Results Gemini 2.0 demonstrated highest sensitivity (68.4%) followed by Claude 3.5 (58.7%) and ChatGPT-4o (19.3%) for fracture detection (p < 0.001). Ideal accuracy rates were 83.3%, 78.7%, and 27.3% respectively. Although ideal accuracy rates exceeded 91% in non-fracture cases, specificity remained low (33.1–36.0%), indicating a high rate of false-positive classifications. Response consistency was very good for ChatGPT-4o (κ = 0.69) and Gemini 2.0 (κ = 0.61), good for Claude 3.5 (κ = 0.44). For Gartland classification, Gemini 2.0 achieved highest accuracy: Type I (83.3%), Type II (62.4%), Type III (68.7%). Conclusion Current LLMs demonstrate limited capability as independent diagnostic tools for pediatric supracondylar humeral fractures. Gemini 2.0’s 68.4% sensitivity indicates these technologies require specialized pediatric training before clinical implementation. However, their potential as assistive tools for triage and assessment warrants further development of pediatric-specific models.

Read PDF

Similar papers

Aug 2026

Large language models demonstrate variable diagnostic performance and a systematic risk of undertriage in surgical triage of feline metacarpal and metatarsal fractures.

Objective To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures. Methods In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and metatarsal fractures were evaluated. Two board-certified veterinary orthopedic surgeons established a reference standard for surgical versus conservative management. Five LLMs (ChatGPT, version 5.2 [OpenAI Inc]; Gemini, version 3 Pro [Alphabet Inc]; Grok, version 4.1 [SpaceXAI]; Qwen, version 3.5 [Alibaba Cloud]; and Claude Sonnet, version 4.5 [Anthropic PBC]; and Claude Sonnet, version 4.5 [Anthropic PBC]) assessed anonymized case summaries using a standardized zero-shot prompt. Model recommendations were compared with the reference standard to calculate accuracy, sensitivity, specificity, and the Cohen κ. Results The reference standard classified 49 cases (67.1%) as surgical and 24 (32.9%) as conservative. ChatGPT achieved the highest performance (accuracy, 84.9%; sensitivity, 79.6%; specificity, 95.8%; κ = 0.69). Other models showed lower performance; Qwen and Claude Sonnet failed to identify any surgical cases (0% sensitivity). A systematic bias toward conservative management was observed across all models, causing high false-negative rates (undertriage). Conclusions LLMs demonstrate highly variable diagnostic performance and a systematic risk of undertriage in feline fracture assessment. While top-performing models approach expert-level agreement, others fail in critical clinical scenarios. Clinical Relevance Independent LLM use for surgical decision-making in feline orthopedics is not recommended due to undertriage risks. These tools require strict clinician oversight for preliminary triage.

S. Okur, Ç. Özkalıpçı, Büşra Baykal et al. · 0 citations
Open access Jul 2026

Artificial intelligence advancements for orthopaedic clinical reasoning: longitudinal assessment of newer models (ChatGPT-5, Grok-3, Gemini 2.5 Flash) compared to clinicians

This descriptive study aimed to longitudinally evaluate the performance of contemporary large language models - ChatGPT-5, Gemini 2.5 Flash, and Grok-3 - on orthopaedic clinical multiple-choice tasks, benchmarked against pooled clinician consensus. A secondary aim was to assess whether recent advances in generative AI translated into improved alignment with clinician consensus compared with previous AI models. A total of 97 multiple-choice clinical cases spanning eight orthopaedic subspecialties were sourced from OrthoBullets and previously benchmarked against aggregated responses from thousands of practising clinicians. Using identical methodology to our 2023 study of ChatGPT-3.5, ChatGPT-4, and Bard, each model was prompted with standardised case stems and response options. The primary outcome was the proportion of AI responses matching the most popular clinician response; secondary analyses assessed agreement within 10% and 20% of clinician consensus, performance on ‘controversial’ (< 25% margin) questions, and inter-model concordance using Cohen’s kappa coefficients. Gemini 2.5 Flash achieved the highest alignment with clinician consensus (69.1%), followed by Grok-3 (66.0%) and ChatGPT-5 (58.8%). None of the LLMs refused to respond to any prompts, representing a reduction from 7.2% from our 2023 study. Subspecialty analysis demonstrated that Gemini 2.5 Flash performed best in Hand and Paediatric domains, while Grok-3 excelled in Reconstruction, Trauma, and ‘controversial’ cases. Inter-model agreement was highest between Grok-3 and Gemini 2.5 Flash (κ = 0.678), indicating improved consistency compared with prior-generation systems. Contemporary LLMs can be promising adjuncts for orthopaedic education by simulating peer reasoning and offering structured explanations in non-critical settings. Despite incremental gains in reasoning capability compared to previous AI models, contemporary LLMs remain unsuitable for independent clinical use. Future research should develop hybrid clinician–AI workflows and longitudinal benchmarks to distinguish true reasoning improvements from memorisation.

Suzen Agharia, Shayan Soroush, Daniel Ameen et al. · 0 citations
Review Open access Aug 2026

Artificial intelligence in clinical decision-making: a comparison of ChatGPT 5.0 and Gemini 3.0 in otologic cases

Large Lingual Models (LLMs) can suggest treatment options based on a patient’s history and symptoms. However, they may lack the ability to deliver fully patient-specific recommendations, which may lead to misdiagnosis or inappropriate treatment decisions. This study compared the diagnostic accuracy, investigation recommendations, and treatment planning consistency of two LLMs, ChatGPT 5.0 and Gemini 3.0, using real-world otologic cases. Fifty retrospective cases with diverse clinical characteristics (symptoms, examination findings, audiometric tests, and radiological sections) were selected. Three experienced otorhinolaryngologists established the final diagnoses as a gold standard. The cases were processed by both LLMs, which were instructed to act as ENT specialists. Two blinded reviewers independently evaluated the responses using the Artificial Intelligence Performance Instrument (AIPI) across four domains: summarization, differential diagnosis, additional examination, and therapy options. Inter-rater reliability was assessed via Cohen’s kappa. Both models demonstrated high performance in managing otologic cases. However, Gemini 3.0 significantly outperformed ChatGPT 5.0 in diagnostic accuracy, investigation/therapy planning, and overall AIPI scores ( p  < 0.05). Subgroup analysis revealed that Gemini 3.0 performed higher AIPI scores in ‘vertigo’ cases, while similar performances were found in ‘otits media’, ‘hearing loss’ and ‘facial paralysis’ groups. Inter-rater agreement for AIPI scores was excellent. To our knowledge, this is the first study evaluating LLMs using real-world otologic data. While both models show potential as clinical decision-support tools, Gemini 3.0 exhibited superior diagnostic performance in real-world otologic cases. Future research should include larger populations and endoscopic imaging to further validate these findings.

Bilge Tuna, Gökhan Tüzemen, Hasan Mutlu · 0 citations
Open access Mar 2026

Can Large Language Models Identify When an Upper Extremity Problem Needs Nonurgent Attention? An Assessment of Multiple LLM Chatbots.

IntroductionSeeking emergency care regarding musculoskeletal sensations is far more prevalent than limb or life-threatening pathophysiology. We studied the ability of an LLM to distinguish between urgent and nonurgent upper extremity symptoms and provide an accurate diagnosis.MethodsFive LLMs (ChatGPT-4, ChatGPT-4o, Co-Pilot, Gemini, and PerplexityChat) were presented with descriptions of seven urgent and seven nonurgent symptom scenarios written below a sixth grade reading level. LLM responses were identified as appropriate if immediate urgent medical attention was recommended after an initial and ongoing inquiry ("What additional information do you need to diagnose my condition?"). Diagnoses provided were identified as correct, partially correct, or incorrect. The analysis was repeated 24 months later with the current LLM versions and results were compared.ResultsLLMs discerned nonurgent conditions with 97% positive predictive value (PPV) and an 89% negative predictive value (NPV) on initial query, which improved to 96% and 97% respectively after ongoing inquiry. Compartment syndrome was misidentified as nonurgent in 80% of scenarios on initial inquiry, although four of five LLMs corrected on continued inquiry. Diagnosis was correct or partially correct for 115 of 150 (82%) on initial inquiry. An updated analysis 24 months later demonstrated marked improvement in LLM ability to identify emergencies with 100% PPV and 95% NPV on initial query and 98% NPV after ongoing inquiry.ConclusionThe finding that LLMs can distinguish urgent from nonurgent upper extremity conditions suggests that artificial intelligence tools could help reduce unnecessary use of high-cost emergency services, allowing those resources to be reserved for patients who require timely care.

Jefferson Hunter, David Ring, Prakash Jayakumar · 0 citations
Open access Aug 2026

How Successful Is Artificial Intelligence in Hand Surgery Questions of the Turkish Orthopedics and Traumatology Board Exams?

Aim: Artificial intelligence (AI) applications are increasingly used in medical education and clinical research. The purpose of this study was to examine the accuracy of various large language models (LLMs) in responding to questions related to hand surgery. This evaluation was based on items derived from the Turkish Orthopedics and Traumatology Board Examination administered between 2010 and 2025.Methods: A total of 220 hand surgery–related multiple-choice questions were extracted from national board examinations administered between 2010 and 2025. Questions were posed to three LLMs (ChatGPT-5.0, Gemini-Pro, and DeepSeek-V3) using both collective and individual questioning approaches across three separate sessions. Model responses were compared with the official correct answers, and success rates were calculated descriptively.Results: All three LLMs achieved satisfactory performance across all years, with success rates ranging from 73.6% to 90.9%. Individually asked questions yielded higher average scores compared with collectively asked questions. Year-by-year analysis demonstrated that all models met or exceeded the examination passing threshold throughout the 16-year period.Conclusion: Current LLMs demonstrate a high level of factual knowledge in hand surgery board examination questions. While these models cannot replace formal medical training or clinical judgment, they may serve as supportive tools in orthopedic education and exam preparation.

Ahmet Acar, A. B. Girgin, Sema Cihan · 0 citations
Review Open access Aug 2026

Artificial Intelligence Answering Patient Questions About Developmental Dysplasia of the Hip: Accuracy, Readability, and Clinical Utility.

INTRODUCTION Artificial intelligence (AI) tools are increasingly used by patients and caregivers seeking medical information. Developmental dysplasia of the hip (DDH) is a common pediatric condition that frequently prompts parental questions regarding screening, diagnosis, and treatment. The purpose of this study was to evaluate the accuracy, completeness, and clarity of responses generated by a Department of Defense-approved AI chatbot (NIPR-GPT) when answering commonly asked DDH-related questions. METHODS Twelve frequently asked patient questions related to DDH were submitted to the NIPR-GPT chatbot. Each question was entered into a new chatbot session to minimize contextual bias. AI-generated responses were independently evaluated by 5 fellowship-trained pediatric orthopedic surgeons. Responses were graded using a 4-point scale assessing accuracy, completeness, and clarity: (1) unsatisfactory (major inaccuracies requiring substantial correction), (2) satisfactory with moderate clarification required, (3) satisfactory with minimal clarification required, and (4) excellent with no clarification required. Inter-rater reliability was assessed using intraclass correlation coefficients (ICC) and weighted kappa statistics, while overall internal consistency among raters was assessed using Krippendorff's alpha. RESULTS All AI-generated responses (100%) were rated either satisfactory or excellent. Six responses (50%) were graded excellent, requiring no clarification, and 6 responses (50%) were graded satisfactory with minimal clarification required. No responses were graded as unsatisfactory or requiring major correction. Single-rater reliability among individual reviewers was poor (ICC [2, 1] = 0.12), reflecting variability in reviewer thresholds. However, reliability improved when scores were averaged across raters, demonstrating moderate agreement (ICC [2, k] = 0.41). Internal consistency among reviewers was Krippendorff's α = 0.163, indicating heterogeneity in reviewer grading but not systematic deficiencies in AI responses. Five fellowship-trained pediatric orthopedic surgeons independently graded all 12 responses. No reviewer graded any response as unsatisfactory or requiring substantial correction. Mean question scores ranged from 2.8 to 4.0 across the 12 questions, with an overall mean score of approximately 3.6. CONCLUSION NIPR-GPT generated generally accurate, clear, and clinically appropriate responses to common DDH-related questions. Half of responses required no clarification, and none required major correction. These findings suggest that AI chatbots may serve as useful adjunct educational tools for patients and caregivers within military healthcare systems. However, variability in expert interpretation and the potential for patient misinterpretation underscore the continued importance of clinician oversight when integrating AI-generated information into patient education.

Michael G Johnston, William J Ferris, Harrison Diaz et al. · 0 citations