LLMs appear to have potential as supplementary information resources for clinicians in restorative dentistry; however, their clinical integration, impact on patient outcomes, and real-world usability remain to be established in future research.
Abstract
Objective
With the advancement of artificial intelligence (AI), large language models (LLMs) have become an alternative source of information in dentistry. These LLMs, which can be trained on large data sets, can answer medical questions and provide references, but they can pose problems in terms of ethics and accurate information. This research aims to evaluate the accuracy of five different LLMs in diagnosis and treatment planning in the field of restorative dentistry.
METHODOLOGY
The 20 most common cases encountered in a restorative dentistry clinic were formulated into questions. The validity of the questions was assessed using the Lawshe Content Validity Index. The questions were posed to five different LLMs: ChatGPT-5, Deepseek V3.2, Claude Sonnet 4.5, Microsoft Copilot, and Google Gemini 3 Flash. Each model was asked to create a diagnosis and treatment plan for each case. The responses were evaluated by 42 restorative dentistry specialists using a Likert Scale. Additionally, the accuracy of the references provided in the responses was evaluated by the article authors. The obtained data were analyzed using the non-parametric Kruskal-Wallis test, and the Dunn multiple comparison test was applied in cases where significant differences were detected.
Results
Statistically significant differences were found between the models for 15 out of 20 questions (p < 0.05). A significant difference was also found in terms of total median scores (p < 0.001), with Google Gemini 3 Flash (median:83) and ChatGPT-5 (median:81) achieving the highest scores. Reference quality was evaluated using a four-category framework. Claude Sonnet 4.5 and Google Gemini 3 Flash demonstrated the highest proportions of accurate and relevant references (85.7% and 84.2%, respectively), while DeepSeek V3.2 exhibited the highest fabrication rate (55.6%).
Conclusions
Based on specialist-evaluated response quality, no model demonstrated consistent and superior performance across all clinical scenarios. LLMs appear to have potential as supplementary information resources for clinicians in restorative dentistry; however, their clinical integration, impact on patient outcomes, and real-world usability remain to be established in future research.
Artificial intelligence (AI) has emerged as a valuable tool in the field of dentistry to assist healthcare providers in diagnosis, optimize treatment outcomes, support research, and improve education. Furthermore, the generative aspects of AI have emerged as a powerful tool for dental professionals to tackle clinical and academic challenges. However, the reliability of AI in domains such as undergraduate training remains unverified.
This study aimed to evaluate the accuracy and consistency of three large language models (LLMs) in answering multiple-choice questions on the undergraduate operative dentistry curriculum to determine their reliability as a supplementary learning resource.
Sixty multiple-choice questions were formulated from the undergraduate operative dentistry curriculum. Each LLM (Claude, Gemini, and ChatGPT) was queried individually nine times over three days. The results were evaluated using a predetermined answer key. The accuracy of the LLMs was compared using a generalized estimating equation (GEE) with question-level clustering analysis. Consistency was assessed at the question level using response consistency, correctness consistency, and Fleiss' kappa. Statistical significance was set at
P
< 0.05.
Gemini demonstrated the highest overall accuracy (98.89%), followed by ChatGPT (98.33%) and Claude (97.78%). No statistically significant difference in accuracy was observed among the three LLMs (GEE, overall Wald
χ
2
(2) = 1.46,
p
= 0.481). All three models showed almost perfect intersession agreement (Fleiss'
κ
= 0.93–0.95), with response and correctness consistency ranging from 88.3% to 91.7% across models.
All three LLMs demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool. However, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.
J. M. Antony, N. Harikrishnan, N. Jayasheelan· Frontiers in Dental Medicine· 0 citations
BACKGROUND
The aim of this study is to compare the performance of three large language models (ChatGPT 5, Microsoft Copilot, and Google Gemini 3), with that of dental students using their responses to multiple-choice questions (MCQs) in restorative dentistry. Accuracy of responses were analyzed across the knowledge and cognitive process dimensions of the revised Bloom's taxonomy (RBT), as well as across subject areas.
METHODS
The restorative dentistry exam questions used in this study were drawn from Turkish Dentistry Specialization Entrance Exam (DUS) administered between 2020 and 2025. The 90 five-option, single-best-answer MCQs were classified according to the RBT to ensure cognitive diversity. Following the exclusion of one exam question which had been annulled by the examination authority, the data analysis was performed on the remaining 89 exam questions. Accuracy of AI models and dental students was compared using Pearson's chi-square test and Monte Carlo-corrected Fisher's exact test. Pairwise comparisons were carried out via Bonferroni-corrected Z-test. The results were presented as frequencies and percentages, and p < 0.050 was considered statistically significant.
RESULTS
Microsoft Copilot and Gemini 3 showed similar performance in answering MCQs, both models achieved higher accuracy than ChatGPT 5 and students (p < 0.001). ChatGPT 5's overall accuracy was found to be significantly higher than that of the students. Accuracy of responses varied according to Bloom's taxonomy levels and subject areas. Microsoft Copilot exhibited over 90% accuracy in all categories of Bloom's knowledge and cognitive process dimensions. At the application level of the cognitive process dimension, all chatbots descriptively achieved 100% accuracy; however, this subgroup difference did not reach statistical significance. Chatbot performance was generally superior to that of students across the subject areas of adhesive dentistry, dentin hypersensitivity, dental caries, tooth whitening, aesthetic restorative procedures, contemporary restorative materials, preventive dentistry, lasers, and saliva.
CONCLUSION
AI-based chatbots demonstrate considerable potential in answering questions about restorative dentistry. At the same time, the differences in performance observed among different models suggest there could be variation in the accuracy values of these systems depending on both the taxonomic level and the subject area of restorative dentistry. The findings support the potential use of chatbots as complementary learning resources in dental education. However, the reliability of such systems should be consistently verified and tested in both an academic and clinical setting under the direct supervision of experts. Moreover, students should be equipped with critical thinking skills to appropriately evaluate and use these systems.
O. Akkurt, Ebru Yılmaz, Nilgün Akgül· BMC Medical Education· 0 citations
PURPOSE
The purpose of this study was to evaluate and compare four AI software programs-ChatGPT-5, Microsoft Copilot, Google Gemini (V 2.5), and OpenEvidence-in generating comprehensive dental treatment plans for minimally destructed and severely mutilated teeth using identical clinical inputs.
MATERIAL AND METHODS
Ten anonymized clinical cases, each consisting of 1 intraoral photograph and 1 corresponding periapical radiograph, were independently submitted to each AI software program using a standardized prompt. A reference standard was established through consensus among four calibrated specialists (one surgically trained prosthodontist, one restorative dentist, one prosthodontist, and one endodontist). AI-generated responses were evaluated using a structured scoring rubric across five domains: diagnostic accuracy, restorability assessment, multidisciplinary integration, treatment sequencing, and extraction appropriateness (score range: 0-10 per case). Agreement among evaluators was assessed using pairwise Cohen's kappa (κ) analysis based on initial independent scoring before consensus discussions.
RESULTS
Substantial inter-evaluator agreement was observed (κ = 0.74; 95% CI: 0.68-0.80), indicating consistent application of the scoring rubric. Mean total scores were highest for OpenEvidence (7.7 ± 1.4) and ChatGPT-5 (7.6 ± 1.1), followed by Google Gemini (5.1 ± 1.2) and Microsoft Copilot (4.7 ± 3.1). OpenEvidence and ChatGPT-5 more frequently incorporated phased treatment sequencing, ferrule assessment, restorability analysis, and multidisciplinary treatment considerations. Microsoft Copilot failed to generate responses in three of the 10 evaluated cases because of content-filtering restrictions.
CONCLUSIONS
AI software programs generated structured and clinically relevant dental treatment proposals; however, meaningful variability existed in diagnostic interpretation, restorability assessment, multidisciplinary integration, and extraction thresholds. OpenEvidence and ChatGPT-5 demonstrated greater agreement with the expert reference standard, although clinically significant diagnostic errors were observed across all evaluated systems. AI-generated treatment plans should therefore be regarded as adjunctive decision-support tools requiring specialist oversight before clinical implementation.
Sohil A Kazim, Bashaer A Alnoman, Razan M Hashim et al.· Journal of Prosthodontics· 0 citations
Objective: The aim of this study was to compare the accuracy, completeness, and readability of information related to whitening in dentistry provided by the large language models (LLM) ChatGPT-5, Gemini 2.5 Pro, DeepSeek v3.2, and Claude Sonnet 4.5. Methods: A total of 25 open-ended questions covering various aspects of dental whitening were prepared and presented to four different LLMs. The responses were recorded and evaluated independently by two restorative dental treatment specialists who were blinded to the source of the responses. Accuracy and completeness were evaluated using 5-point and 3-point Likert scales respectively. Readability was evaluated using the Flesch Reading Ease Score (FRES), the Flesch–Kincaid Grade Level (FKGL), and the Simple Measure of Gobbledygook (SMOG) index. In the statistical analyses, the Intraclass Correlation Coefficient (ICC), Shapiro-Wilk test, Kruskal-Wallis test, Dunn-Bonferroni test, and Spearman correlation analysis were used. Results: A statistically significant difference was identified among the artificial intelligence models regarding the accuracy of their responses to open-ended questions related to dental whitening (p<0.05). The accuracy score of DeepSeek v.3.2 (4.52±0.77) was determined to be statistically significntly higher than that of ChatGPT-5 (3.88±1.01). The FRES points for readability were found to be statistically significantly higher for DeepSeek v3.2 (42.02±7.63) compared to the mean points obtained for ChatGPT-5 (29.09±9.07), Gemini 2.5 Pro (33.92±9.24), and Claude Sonnet 4.5 (21.65±9.29). Conclusion: DeepSeek v3.2 exhibited better performance than ChatGPT-5 in terms of accuracy. No statistically significant differences were identified among the models regarding the completeness scores of responses to open-ended questions. However, the FRES results indicated that DeepSeek v3.2 generated texts with higher readability compared to the other chatbots. Overall, these findings suggest that DeepSeek v3.2 demonstrated superior performance in terms of both accuracy and readability.
INTRODUCTION
Large language models (LLMs) are increasingly utilized for medical and dental information retrieval, yet their ability to interpret authentic, patient-style inquiries remains insufficiently investigated. This study compared the performance of ChatGPT, Claude, and Gemini in responding to patient-oriented queries related to periodontal and peri-implant diseases.
MATERIALS AND METHODS
Unlike traditional investigations using expert-generated questions, this study employed 40 realistic, patient-oriented queries designed to simulate the post-examination cognitive state, blending colloquial language with partially retained clinical jargon. Each query was submitted to GPT-4o, Claude Sonnet 5 and Gemini 2.5 Pro generating 120 responses. Three blinded periodontists independently evaluated scientific accuracy, completeness, clinical safety, and overall quality using a 5-point Likert scale. Automated text analysis assessed readability metrics (Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index) and linguistic characteristics. Statistical protocols included Friedman, Bonferroni-adjusted Wilcoxon signed-rank, and Model Dominance analyses.
RESULTS
Significant performance differences were observed among the models across all expert-rated domains (all p < 0.001). Gemini achieved the highest expert ratings for scientific accuracy (4.77 ± 0.22), clinical safety (4.94 ± 0.15), and overall quality (4.85 ± 0.20), and was identified as the most frequently top-ranked platform via dominance analysis. Conversely, Claude performed significantly better regarding response completeness (4.77 ± 0.22) and demonstrated the most favorable overall readability profile, yielding the lowest Flesch-Kincaid Grade Level (7.54 ± 1.27). GPT-4o consistently received the lowest expert ratings across all evaluated domains.
DISCUSSION
While all evaluated LLMs generated high-quality responses to realistic periodontal queries, their functional strengths were highly multidimensional. Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations. These findings support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.
Ramazan Ağırağaç, Vedat Yüksekkaya· Journal of Stomatology Oral...· 0 citations
INTRODUCTION
The use of artificial intelligence (AI) in orthodontic practice is increasing rapidly; however, there is a notable lack of research evaluating the accuracy of large language models (LLMs) in educating patients about orthodontic retainers and related guidelines.
MATERIALS AND METHODS
This study utilized a cross-sectional, repeated‑measures comparative evaluation design after receiving exemption from the institutional ethics committee. A set of 110 questions related to orthodontic retainers was compiled from previous articles addressing concerns about retainers and approved by a panel of three orthodontists. These questions were submitted to large language models (LLMs), including ChatGPT, Copilot, DeepSeek and Google Gemini. The responses were then reviewed by six independent orthodontists, who rated them using a modified five-point Likert scale.
RESULTS
The overall accuracy revealed that 68.6% of responses scored 4, while 15.3% achieved a perfect score of 5. Among the LLMs, Gemini ranked first with 96.8%, closely followed by ChatGPT at 95.6%, indicating comparable high‑level performance between these models, while DeepSeek (76.9%) and Copilot (66.2%) demonstrated comparatively lower accuracy. Gemini produced a higher proportion of perfect scores, whereas ChatGPT consistently achieved strong ratings. The mean ratings across six raters demonstrated strong reliability (ICC = 0.81), reflecting expert agreement.
CONCLUSIONS
Findings suggest that AI models such as ChatGPT and Gemini can generate patient‑directed orthodontic retainer information with high informational accuracy under controlled evaluation conditions. However, specialist oversight remains essential to ensure clinical applicability. Future research using larger and more diverse datasets is needed to assess broader educational and communication‑related outcomes.
Muhammad Mughni, M. Ilyas, Syed Ali Naqi Gilani et al.· BMC Oral Health· 0 citations