Language-dependent performance variation in large language models for dental trauma management: a comparative evaluation of ChatGPT-5.2, Gemini 3.0, and Claude 4.5 Sonnet.
This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management.
Abstract
Background
Large language models (LLMs) are increasingly evaluated for medical question answering and clinical information tasks, yet the impact of query language on their performance in specialized domains such as dental traumatology remains insufficiently studied. The primary objective was to evaluate whether query language (English vs. Turkish) affects LLM performance in a controlled scenario-based assessment of dental trauma management. Secondary objectives were to compare overall performance across three LLMs and to examine whether language effects are uniform across models or model-specific.
Methods
Twenty-seven clinical scenarios covering 13 dental trauma categories were presented to ChatGPT 5.2, Gemini 3.0, and Claude 4.5 Sonnet in both English and Turkish, generating 162 responses. Two blinded endodontists independently evaluated responses using a standardized rubric assessing accuracy (40%), completeness (35%), and safety (25%) against IADT 2020 Guidelines. Inter-rater reliability was assessed using intraclass correlation coefficient (ICC). Language effects were analyzed using Wilcoxon signed-rank tests; model comparisons employed Kruskal-Wallis and Mann-Whitney U tests with Bonferroni correction.
Results
Inter-rater reliability ranged from moderate to good across evaluation dimensions (ICC: 0.738-0.836). ChatGPT showed the strongest language effect with 9.14% higher performance in English (p < 0.001, r = 0.874). Gemini showed moderate English advantage (5.69%, p = 0.003, r = 0.572). Claude exhibited language independence with virtually identical performance in both languages (-0.02%, p = 0.220). In English, significant model differences emerged (H = 22.31, p < 0.001); however, model performance converged in Turkish (H = 2.89, p = 0.236).
Conclusions
This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management. ChatGPT 5.2 achieved the highest performance in English but exhibited the most pronounced Turkish-language degradation, including substantial safety score decline. Gemini 3.0 showed an intermediate pattern with moderate English advantage. Claude 4.5 Sonnet demonstrated language-independent performance across all evaluated dimensions. These findings are based on a standardized scenario-based assessment and should not be extrapolated to real clinical environments or patient care settings.
Large language models (LLMs) are increasingly used to answer medical questions; however, their performance may vary depending on task characteristics. This study evaluated the performance of multiple versions of two widely used LLM families on oral and maxillofacial radiology (OMFR) questions from the Turkish Dental Specialty Examination (DUS) across three assessment phases and examined the influence of cognitive complexity, model family, evaluation phase, and content domain on response accuracy. A comparative repeated-evaluation design was used. A total of 123 text-based OMFR questions from DUS examinations (2012-2021) were submitted to two widely used LLM families (ChatGPT and DeepSeek) across three evaluation phases (May 2025, August 2025, and February 2026). Questions were categorized by content domain and Bloom cognitive level (low vs. high). Model responses were evaluated using official answer keys, and generalized estimating equations (GEE) were applied to account for repeated measurements. A total of 1230 model responses were analyzed, yielding an overall accuracy of 83.7%. Agreement between repeated runs was substantial to almost perfect (κ = 0.689-0.912). Cognitive complexity emerged as the strongest determinant of performance, with low-level questions significantly more likely to be answered correctly than high-level questions (OR = 6.15, p = 0.003). Content domain was also associated with accuracy (p = 0.028), whereas no statistically significant associations were observed for model family or evaluation phase. LLMs demonstrated high accuracy in answering OMFR examination questions; however, performance was more strongly associated with cognitive complexity than with model family or evaluation phase.
BACKGROUND
Large language models (LLMs), including ChatGPT, Gemini and DeepSeek, are increasingly explored as supportive tools in health professions education. However, their educational utility across different knowledge domains and question formats, particularly those involving visual content, remains insufficiently understood.
OBJECTIVE
This study aimed to evaluate the educational potential of three advanced LLMs in dental training by analysing their performance on a national specialty examination over a 10-year period, with particular emphasis on domain-specific accuracy and differences between text-based and visual questions.
METHODS
A total of 1560 multiple-choice questions from 13 administrations of the Turkish Dental Specialty Examination (DUS) between 2012 and 2021 were included. All questions were translated into English and categorized into nine dental specialties. Each question was individually entered into ChatGPT-4.0, Gemini Advanced and DeepSeek in isolated sessions. Model responses were compared with official answer keys, and accuracy rates were analysed across years, specialties and question types. Statistical analyses included chi-square tests, one-way ANOVA and Kruskal-Wallis tests.
RESULTS
ChatGPT-4.0 achieved the highest overall accuracy (86.93%), followed by Gemini (82.86%) and DeepSeek (82.34%). Performance varied across specialties, with higher accuracy observed in Basic Sciences and Oral Surgery, and lower performance in Orthodontics and Endodontics. A substantial decrease in accuracy was observed for visual questions (ChatGPT: 52.38%; Gemini: 45.24%; DeepSeek: 2.38%) compared to text-based items (all models > 83%). ChatGPT demonstrated more stable performance across years, whereas Gemini and DeepSeek showed greater variability.
CONCLUSIONS
LLMs demonstrate strong potential as supportive tools in dental education, particularly for text-based knowledge assessment. However, their limited performance in visual question contexts highlights an important constraint for their integration into image-dependent domains such as dental diagnostics. These findings underscore the need for cautious and context-aware implementation of AI tools in dental curricula and assessment practices.
Sedef Ayşe Taşyapan, Didem Özer· European journal of dental e...· 0 citations
ObjectiveThis study evaluates the performance of large language models (LLMs)-ChatGPT-4.0, Gemini 2.0 Pro, o3-mini, Doctor GPT and DeepSeek-V3-in a national orthopaedic proficiency examination and explores their implications for health informatics and medical education. The responses of these models were analysed to assess accuracy rates and differences between models.MethodA total of 100 multiple-choice questions from the 2024 TOTEK examination were administered to each AI model under identical conditions. Correct and incorrect responses were recorded, and differences in performance were evaluated using chi-square testing and frequency analysis. Question categories were also compared to identify domain-specific variations.Resultso3-mini achieved the highest accuracy rate (79%), while Gemini 2.0 showed the lowest (68%); all models exceeded the 60% pass threshold. A statistically significant difference between models was identified in the Surgical Procedures category, in which Gemini 2.0 answered fewer questions correctly (23/36) than the other models (30-32/36) (χ2 = 9.87, df = 4, p = 0.043). No significant differences were observed in the remaining categories (all p > 0.05), and the overall difference in accuracy between models did not reach statistical significance (χ2 = 4.01, df = 4, p = 0.405). Clinical decision-making and visual content-based questions were the most challenging for all models.ConclusionAI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.
Bünyamin Arı· Health Informatics Journal· 0 citations
INTRODUCTION
Large language models (LLMs) are increasingly used as sources of information across many fields, including healthcare. As patients turn to these models for health-related queries, evaluating the accuracy and reliability of their responses is essential. Peri-acetabular osteotomy (PAO) is performed on younger patients - a group more likely to use digital tools like LLMs for health information. This study assesses the accuracy and readability performance of two leading LLMs, ChatGPT and Google Gemini, in addressing common patient questions on PAO.
METHODS
A panel of fellowship-trained PAO surgeons created ten commonly asked patient questions based on real-world experience. Responses from each LLM were assessed by the same three surgeons, blinded to response origin, using a 5-point Likert scale to evaluate clarity, accuracy, and completeness. Readability was measured with Flesch-Kincaid Reading scores.
RESULTS
ChatGPT outperformed Gemini with an average score of 4.17 vs 3.13 (t = -3.08, p = 0.006). ChatGPT's responses were often rated higher for completeness and clarity, particularly in areas needing detailed explanation, and usually required minimal clarification. Gemini sometimes lacked specificity or included minor inaccuracies that reduced its perceived reliability. Both LLMs produced responses with similar "difficult" Flesch Reading Ease scores.
CONCLUSIONS
There may be significant differences in how effectively LLMs support patients with surgical queries. ChatGPT more consistently met expert standards for clarity and thoroughness. As LLM usage expands, ChatGPT may aid patient education on hip surgery, supporting consultations, informed decisions and postoperative guidance.
TP Davis, B. Guevel, K. Logishetty et al.· Annals of the Royal College...· 0 citations
INTRODUCTION
Large language models (LLMs) are increasingly utilized for medical and dental information retrieval, yet their ability to interpret authentic, patient-style inquiries remains insufficiently investigated. This study compared the performance of ChatGPT, Claude, and Gemini in responding to patient-oriented queries related to periodontal and peri-implant diseases.
MATERIALS AND METHODS
Unlike traditional investigations using expert-generated questions, this study employed 40 realistic, patient-oriented queries designed to simulate the post-examination cognitive state, blending colloquial language with partially retained clinical jargon. Each query was submitted to GPT-4o, Claude Sonnet 5 and Gemini 2.5 Pro generating 120 responses. Three blinded periodontists independently evaluated scientific accuracy, completeness, clinical safety, and overall quality using a 5-point Likert scale. Automated text analysis assessed readability metrics (Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index) and linguistic characteristics. Statistical protocols included Friedman, Bonferroni-adjusted Wilcoxon signed-rank, and Model Dominance analyses.
RESULTS
Significant performance differences were observed among the models across all expert-rated domains (all p < 0.001). Gemini achieved the highest expert ratings for scientific accuracy (4.77 ± 0.22), clinical safety (4.94 ± 0.15), and overall quality (4.85 ± 0.20), and was identified as the most frequently top-ranked platform via dominance analysis. Conversely, Claude performed significantly better regarding response completeness (4.77 ± 0.22) and demonstrated the most favorable overall readability profile, yielding the lowest Flesch-Kincaid Grade Level (7.54 ± 1.27). GPT-4o consistently received the lowest expert ratings across all evaluated domains.
DISCUSSION
While all evaluated LLMs generated high-quality responses to realistic periodontal queries, their functional strengths were highly multidimensional. Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations. These findings support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.
Ramazan Ağırağaç, Vedat Yüksekkaya· Journal of Stomatology Oral...· 0 citations