Skip to content
Open access

ARE AI MODELS READY FOR CLINICAL DECISION-MAKING IN DENTAL BLEACHING? AN EVIDENCE-BASED EVALUATION OF ACCURACY, COMPLETENESS, AND READABILITY OF CHATGPT-5, GEMINI 2.5 PRO, DEEPSEEK V3.2, AND CLAUDE SONNET 4.5

Jul 2026 · Acta Medica Nicomedia · 0 citations · 36 references

Abstract

Objective: The aim of this study was to compare the accuracy, completeness, and readability of information related to whitening in dentistry provided by the large language models (LLM) ChatGPT-5, Gemini 2.5 Pro, DeepSeek v3.2, and Claude Sonnet 4.5. Methods: A total of 25 open-ended questions covering various aspects of dental whitening were prepared and presented to four different LLMs. The responses were recorded and evaluated independently by two restorative dental treatment specialists who were blinded to the source of the responses. Accuracy and completeness were evaluated using 5-point and 3-point Likert scales respectively. Readability was evaluated using the Flesch Reading Ease Score (FRES), the Flesch–Kincaid Grade Level (FKGL), and the Simple Measure of Gobbledygook (SMOG) index. In the statistical analyses, the Intraclass Correlation Coefficient (ICC), Shapiro-Wilk test, Kruskal-Wallis test, Dunn-Bonferroni test, and Spearman correlation analysis were used. Results: A statistically significant difference was identified among the artificial intelligence models regarding the accuracy of their responses to open-ended questions related to dental whitening (p<0.05). The accuracy score of DeepSeek v.3.2 (4.52±0.77) was determined to be statistically significntly higher than that of ChatGPT-5 (3.88±1.01). The FRES points for readability were found to be statistically significantly higher for DeepSeek v3.2 (42.02±7.63) compared to the mean points obtained for ChatGPT-5 (29.09±9.07), Gemini 2.5 Pro (33.92±9.24), and Claude Sonnet 4.5 (21.65±9.29). Conclusion: DeepSeek v3.2 exhibited better performance than ChatGPT-5 in terms of accuracy. No statistically significant differences were identified among the models regarding the completeness scores of responses to open-ended questions. However, the FRES results indicated that DeepSeek v3.2 generated texts with higher readability compared to the other chatbots. Overall, these findings suggest that DeepSeek v3.2 demonstrated superior performance in terms of both accuracy and readability.

Read PDF