Diagnostic Potential of Artificial Intelligence for Detecting Vertical Root Fractures on Dental Imaging: A Systematic Review and Diagnostic Meta-analysis
Abstract
Vertical root fracture (VRF) remains a diagnostic challenge because radiographic signs are often subtle, non-specific and influenced by imaging modality, fracture characteristics, artefacts and observer interpretation. This systematic review and meta-analysis assessed the accuracy of artificial intelligence (AI)-based models for VRF detection on cone-beam computed tomography (CBCT), periapical radiographs and panoramic radiographs and evaluated the certainty of evidence. This review followed Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy Studies recommendations and was prospectively registered in International Prospective Register of Systematic Reviews. Studies evaluating AI-based models for VRF detection on dental imaging were included. Sensitivity and specificity were pooled using bivariate random-effects models when complete 2 × 2 contingency tables were available. Summary receiver operating characteristic curves were generated. Risk of bias was assessed using Quality Assessment of Diagnostic Accuracy Studies-2 and certainty of evidence was evaluated using the Grading of Recommendations Assessment, Development and Evaluation approach adapted for diagnostic test accuracy. Post-hoc sensitivity analyses retained one representative dataset per study to assess the influence of multiple model-specific datasets. Six studies were included, of which 4 contributed to quantitative synthesis. The CBCT-based AI models showed high diagnostic performance under predominantly controlled or partially controlled conditions, with a pooled sensitivity of 85.3% and specificity of 89.0% (area under the curve [AUC]=0.921). The CBCT convolutional neural network–only subgroup showed sensitivity of 88.3% and specificity of 87.6% (AUC=0.929). Periapical radiograph-based models showed high sensitivity but limited specificity, with pooled sensitivity of 86.9% and specificity of 61.2% (AUC=0.708). The probabilistic neural network-only periapical subgroup improved sensitivity, but specificity remained limited. Sensitivity analyses confirmed preserved CBCT performance and persistently low periapical specificity. Panoramic radiography was assessed qualitatively because complete 2 × 2 data were unavailable. Certainty of evidence was moderate for CBCT and low for periapical radiographs. AI-based models show promising diagnostic potential for VRF detection, particularly on CBCT. However, current evidence mainly reflects performance under controlled or partially controlled conditions and should not be interpreted as definitive real-world diagnostic effectiveness. AI should currently be considered an adjunctive decision-support tool rather than an autonomous diagnostic method, especially for 2-dimensional imaging modalities. Prospective, multicentre clinical validation studies are needed.