Publicly available LLMs can pose potential safety risks to dental patients when used without guardrails and language-specific model training and hallucination reduction strategies, such as retrieval-augmented generation, are recommended.
Abstract
Background
Patients increasingly turn to the Internet and, more recently, to Large Language Models (LLMs) to seek answers to dental questions. However, the safety of LLM-generated dental advice for patients has not been systematically assessed.
Methods
Thirty dental questions spanning six disciplines were developed and posed in Polish and English to six LLMs. Each question was answered five times per model per language, yielding 2,700 unique LLM-generated answers, each independently evaluated by two dentists. A binary text classifier combining Qwen3-Embedding-0.6B word embeddings with a Support Vector Machine (SVM) was trained on 1,680 labeled answers and evaluated on 420 held-out answers.
Results
Substantial inter-rater agreement on harmfulness was observed (Cohen's [Formula: see text]). Without a system prompt, 56% of outputs from gpt-4o, gpt-4o-mini, and llama-3.3-70b were classified as harmful (pooling both languages; a question was counted as harmful if at least one of its five generated answers was labeled harmful by at least one annotator). A safety-oriented system prompt reduced this rate by 26 percentage points to 30%. llama-3.3-70b showed markedly higher harmfulness in Polish than English ([Formula: see text], [Formula: see text]). The trained classifier showed a tendency to mark non-harmful answers as harmful, though performance was lower for Polish than English outputs.
Conclusions
Publicly available LLMs can pose potential safety risks to dental patients when used without guardrails. System prompts significantly mitigate harmful outputs. Language-specific model training and hallucination reduction strategies, such as retrieval-augmented generation, are recommended as future directions.
Large language models (LLMs) are increasingly used to answer medical questions; however, their performance may vary depending on task characteristics. This study evaluated the performance of multiple versions of two widely used LLM families on oral and maxillofacial radiology (OMFR) questions from the Turkish Dental Specialty Examination (DUS) across three assessment phases and examined the influence of cognitive complexity, model family, evaluation phase, and content domain on response accuracy. A comparative repeated-evaluation design was used. A total of 123 text-based OMFR questions from DUS examinations (2012-2021) were submitted to two widely used LLM families (ChatGPT and DeepSeek) across three evaluation phases (May 2025, August 2025, and February 2026). Questions were categorized by content domain and Bloom cognitive level (low vs. high). Model responses were evaluated using official answer keys, and generalized estimating equations (GEE) were applied to account for repeated measurements. A total of 1230 model responses were analyzed, yielding an overall accuracy of 83.7%. Agreement between repeated runs was substantial to almost perfect (κ = 0.689-0.912). Cognitive complexity emerged as the strongest determinant of performance, with low-level questions significantly more likely to be answered correctly than high-level questions (OR = 6.15, p = 0.003). Content domain was also associated with accuracy (p = 0.028), whereas no statistically significant associations were observed for model family or evaluation phase. LLMs demonstrated high accuracy in answering OMFR examination questions; however, performance was more strongly associated with cognitive complexity than with model family or evaluation phase.
INTRODUCTION
Large language models (LLMs) are increasingly utilized for medical and dental information retrieval, yet their ability to interpret authentic, patient-style inquiries remains insufficiently investigated. This study compared the performance of ChatGPT, Claude, and Gemini in responding to patient-oriented queries related to periodontal and peri-implant diseases.
MATERIALS AND METHODS
Unlike traditional investigations using expert-generated questions, this study employed 40 realistic, patient-oriented queries designed to simulate the post-examination cognitive state, blending colloquial language with partially retained clinical jargon. Each query was submitted to GPT-4o, Claude Sonnet 5 and Gemini 2.5 Pro generating 120 responses. Three blinded periodontists independently evaluated scientific accuracy, completeness, clinical safety, and overall quality using a 5-point Likert scale. Automated text analysis assessed readability metrics (Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index) and linguistic characteristics. Statistical protocols included Friedman, Bonferroni-adjusted Wilcoxon signed-rank, and Model Dominance analyses.
RESULTS
Significant performance differences were observed among the models across all expert-rated domains (all p < 0.001). Gemini achieved the highest expert ratings for scientific accuracy (4.77 ± 0.22), clinical safety (4.94 ± 0.15), and overall quality (4.85 ± 0.20), and was identified as the most frequently top-ranked platform via dominance analysis. Conversely, Claude performed significantly better regarding response completeness (4.77 ± 0.22) and demonstrated the most favorable overall readability profile, yielding the lowest Flesch-Kincaid Grade Level (7.54 ± 1.27). GPT-4o consistently received the lowest expert ratings across all evaluated domains.
DISCUSSION
While all evaluated LLMs generated high-quality responses to realistic periodontal queries, their functional strengths were highly multidimensional. Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations. These findings support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.
Ramazan Ağırağaç, Vedat Yüksekkaya· Journal of Stomatology Oral...· 0 citations
This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management.
Background As society is increasingly depending on large language models (LLMs) for health-related questions, it is essential to objectively evaluate the quality and accessibility of the oral cancer information they provide. Although LLMs occupy a growing space in digital health communication, it remains unknown whether the information they generate is both reliable and easy to read. Objective The purpose of this study was to evaluate the reliability and readability, respectively, of the responses generated by four mainstream LLMs (ChatGPT, Gemini, Perplexity, and DeepSeek) to questions related to oral cancer. Specifically, the present authors aimed to evaluate the reliability of responses to common oral cancer questions and assess whether the readability of responses meets established expectations. Methods Twenty-two commonly asked, patient-orientated oral cancer–related questions were developed through two predefined phases: Google Trends analysis and expert consultation with specialists in oral oncology. Each question was entered as an independent single-turn prompt into four LLMs: ChatGPT-5, Gemini 2.5, Perplexity Pro, and DeepSeek v3.2. The primary outcome was information reliability and quality, assessed using four standardized instruments: the DISCERN questionnaire, the Ensuring Quality Information for Patients (EQIP) tool, the Journal of the American Medical Association (JAMA) benchmark criteria, and the Global Quality Scale (GQS). The secondary outcome was readability, assessed using six established indices: the Automated Readability Index, Flesch Reading Ease Score, Gunning Fog Index, Flesch-Kincaid Grade Level, Coleman-Liau Index, and Simple Measure of Gobbledygook. Results Significant differences were observed among the four LLMs in DISCERN, EQIP, and JAMA scores (all P < 0.001), whereas no significant difference was found in GQS scores (P = 0.440). Perplexity Pro achieved the highest mean DISCERN score (46.36 ± 4.70), EQIP score (85.00 ± 0.00), GQS score (4.05 ± 0.58), and JAMA score (1.00 ± 0.00). However, all models produced responses above the recommended sixth-grade readability level. The mean FKGL scores ranged from 12.65 ± 3.07 for ChatGPT-5 to 15.65 ± 3.36 for Perplexity Pro, and the mean FRES scores ranged from 37.50 ± 14.27 for Perplexity Pro to 52.59 ± 12.47 for Gemini 2.5. Conclusion Current LLMs may support oral cancer patient education, but their use remains limited by variable information quality, insufficient transparency, and poor readability. Although Perplexity Pro performed better on several reliability-related metrics, no model showed consistently high performance across all dimensions or met recommended readability standards. Future LLM-based patient education tools should prioritise verifiable sourcing, guideline-based accuracy, risk communication, and plain-language adaptation.
Bo Zhang, Weidi Shi, Ying Zhang· Oral Health & Preventive Den...· 0 citations
Artificial intelligence (AI) has emerged as a valuable tool in the field of dentistry to assist healthcare providers in diagnosis, optimize treatment outcomes, support research, and improve education. Furthermore, the generative aspects of AI have emerged as a powerful tool for dental professionals to tackle clinical and academic challenges. However, the reliability of AI in domains such as undergraduate training remains unverified.
This study aimed to evaluate the accuracy and consistency of three large language models (LLMs) in answering multiple-choice questions on the undergraduate operative dentistry curriculum to determine their reliability as a supplementary learning resource.
Sixty multiple-choice questions were formulated from the undergraduate operative dentistry curriculum. Each LLM (Claude, Gemini, and ChatGPT) was queried individually nine times over three days. The results were evaluated using a predetermined answer key. The accuracy of the LLMs was compared using a generalized estimating equation (GEE) with question-level clustering analysis. Consistency was assessed at the question level using response consistency, correctness consistency, and Fleiss' kappa. Statistical significance was set at
P
< 0.05.
Gemini demonstrated the highest overall accuracy (98.89%), followed by ChatGPT (98.33%) and Claude (97.78%). No statistically significant difference in accuracy was observed among the three LLMs (GEE, overall Wald
χ
2
(2) = 1.46,
p
= 0.481). All three models showed almost perfect intersession agreement (Fleiss'
κ
= 0.93–0.95), with response and correctness consistency ranging from 88.3% to 91.7% across models.
All three LLMs demonstrated high accuracy on operative dentistry MCQs, suggesting their potential utility as a supplementary educational tool. However, the residual error rates in the answers warrant the cautious integration of LLMs into dental education.
J. M. Antony, N. Harikrishnan, N. Jayasheelan· Frontiers in Dental Medicine· 0 citations
To compare reasoning vs. conventional large language models (LLMs) in generating answers with guideline-aligned explanations for Breast Imaging Reporting and Data System (BI-RADS) educational questions.
In this prospective study performed from February 6 to 12, 2025, 49 English-Chinese question pairs were extracted from BI-RADS Atlas Fifth Edition. Two reasoning LLMs (ChatGPT-o1, Deepseek-R1) and six conventional LLMs (Gemini2.0-Flash, Deepseek-V3, ChatGPT-4o, ChatGPT-3.5, Qwen-2.5, and WenXinYiYan-3.5) generated answers and explanations to the questions through structured prompts. Three radiologists specialized in breast imaging independently evaluated responses using a 5-point Likert scale, with reference to standard answers.
The reasoning LLMs significantly outperformed conventional models (median [interquartile range (IQR)]: 3.7 [2.7–4.0] vs. 2.7 [2.0–3.7],
P
< 0.001), with ChatGPT-o1 and Deepseek-R1 demonstrating peak performance. Both categories of LLMs exhibited significant score reductions in handling questions with multifaceted clinical scenarios (reasoning models: median 4.0 [2.7–4.3] vs. 2.7 [2.3–2.7],
Δ
median = −1.3,
P
< 0.001; conventional models: 2.7 [2.0–3.7] vs. 2.3 [2.0–2.7],
Δ
median = −0.4,
P
< 0.001). While question language showed no significant impact on reasoning LLMs (ChatGPT-o1 and Deepseek-R1), it affected some conventional models (ChatGPT-3.5, Deepseek-V3 and Gemini2.0-Flash). LLMs performance remained independent of question section and question type.
Reasoning LLMs show significant potential for BI-RADS guideline explanation and education, but require specific optimization for complex clinical scenario instruction.
Yu-Xia Tang, Yu-Ting Liu, Meng-Xuan Liu et al.· Frontiers in Medicine· 0 citations