INTRODUCTION
Accurate localization of the mandibular canal in Cone-Beam Computed Tomography (CBCT) images is critical for preventing iatrogenic nerve injury during maxillofacial surgery and dental implant procedures. This systematic review and meta-analysis aimed to evaluate the diagnostic performance, anatomical localization accuracy, and time efficiency of deep learning-based artificial intelligence (AI) systems in automated mandibular canal segmentation compared to traditional manual expert annotations.
MATERIALS AND METHODS
A comprehensive literature search was conducted across PubMed, Scopus, Web of Science, IEEE Xplore, and Embase databases in accordance with PRISMA guidelines. Studies evaluating the performance of AI models for mandibular canal detection on CBCT scans using expert annotations as the reference standard were included. The primary outcome measure was the Dice Similarity Coefficient (DSC), while secondary outcomes included Average Symmetric Surface Distance (ASSD) and processing time. Statistical analyses were performed using a random-effects model.
RESULTS
A total of 38 unique studies comprising over 8,420 CBCT volumes were included in the quantitative synthesis. The pooled DSC for AI-driven segmentation was calculated as 0.82 (95% CI: 0.79-0.85). Subgroup analyses revealed that transformer-based architectures (DSC: 0.89) demonstrated significantly superior performance compared to traditional convolutional neural networks (CNNs). The pooled ASSD exhibited a high anatomical accuracy of 0.42 mm (95% CI: 0.38-0.47), which is close to voxel dimensions. Furthermore, the autonomous segmentation process was completed in an average of 32 seconds, whereas manual expert annotation took 600 seconds (p < 0.001), confirming an 18.7-fold timesaving in the clinical workflow.
DISCUSSION
Deep learning algorithms provide highly accurate, reproducible, and time-efficient results at a human-expert level in the automated segmentation of the mandibular canal on CBCT images. The integration of these AI systems into clinical protocols has the potential to enhance surgical safety and standardize preoperative planning processes in dental implantology.
Ramazan Ağırağaç· Journal of Stomatology Oral...· 0 citations
INTRODUCTION
Large language models (LLMs) are increasingly utilized for medical and dental information retrieval, yet their ability to interpret authentic, patient-style inquiries remains insufficiently investigated. This study compared the performance of ChatGPT, Claude, and Gemini in responding to patient-oriented queries related to periodontal and peri-implant diseases.
MATERIALS AND METHODS
Unlike traditional investigations using expert-generated questions, this study employed 40 realistic, patient-oriented queries designed to simulate the post-examination cognitive state, blending colloquial language with partially retained clinical jargon. Each query was submitted to GPT-4o, Claude Sonnet 5 and Gemini 2.5 Pro generating 120 responses. Three blinded periodontists independently evaluated scientific accuracy, completeness, clinical safety, and overall quality using a 5-point Likert scale. Automated text analysis assessed readability metrics (Flesch Reading Ease, Flesch-Kincaid Grade Level, Gunning Fog Index) and linguistic characteristics. Statistical protocols included Friedman, Bonferroni-adjusted Wilcoxon signed-rank, and Model Dominance analyses.
RESULTS
Significant performance differences were observed among the models across all expert-rated domains (all p < 0.001). Gemini achieved the highest expert ratings for scientific accuracy (4.77 ± 0.22), clinical safety (4.94 ± 0.15), and overall quality (4.85 ± 0.20), and was identified as the most frequently top-ranked platform via dominance analysis. Conversely, Claude performed significantly better regarding response completeness (4.77 ± 0.22) and demonstrated the most favorable overall readability profile, yielding the lowest Flesch-Kincaid Grade Level (7.54 ± 1.27). GPT-4o consistently received the lowest expert ratings across all evaluated domains.
DISCUSSION
While all evaluated LLMs generated high-quality responses to realistic periodontal queries, their functional strengths were highly multidimensional. Gemini demonstrated superior clinical precision and safety, whereas Claude provided more comprehensive and readable explanations. These findings support the integration of LLMs as pragmatic, high-ecological-validity complementary tools for patient education, while emphasizing the persistent necessity for professional clinical oversight.
Ramazan Ağırağaç, Vedat Yüksekkaya· Journal of Stomatology Oral...· 0 citations