Accuracy and response repeatability of three large language models on undergraduate operative dentistry multiple-choice questions
Abstract
Artificial intelligence (AI) has emerged as a valuable tool in the field of dentistry to assist healthcare providers in diagnosis, optimize treatment outcomes, support research, and improve education. Furthermore, the generative aspects of AI have emerged as a powerful tool for dental professionals to tackle clinical and academic challenges. However, the reliability of AI in domains such as undergraduate training remains unverified. This study aimed to evaluate the accuracy and consistency of three large language models (LLMs) in answering multiple-choice questions on the undergraduate operative dentistry curriculum to determine their reliability as a supplementary learning resource. Sixty multiple-choice questions were formulated from the undergraduate operative dentistry curriculum. Each LLM (Claude, Gemini, and ChatGPT) was queried individually nine times over three days. The results were evaluated using a predetermined answer key. The accuracy of the LLMs was compared using a generalized estimating equation (GEE) with question-level clustering analysis. Consistency was assessed at the question level using response consistency, correctness consistency, and Fleiss' kappa. Statistical significance was set at P < 0.05. Gemini demonstrated the highest overall accuracy (98.89%), followed by ChatGPT (98.33%) and Claude (97.78%). No statistically significant difference in accuracy was observed among the three LLMs (GEE, overall Wald χ 2 (2) = 1.46, p = 0.481). All three models showed almost perfect intersession agreement (Fleiss' κ = 0.93–0.95), with response and correctness consistency ranging from 88.3% to 91.7% across models. All three LLMs demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool. However, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.