Aug 2026· Health Science Reports· Vol 9· 0 citations
Medicine
TL;DR
AI-based models, particularly CNN-based DL architectures, demonstrate clinically relevant diagnostic performance with high sensitivity, specificity, and diagnostic odds ratios, supporting their potential role as adjunctive tools in CBCT-based differentiation of OKCs from other odontogenic lesions.
Abstract
ABSTRACT Background and Aims Differentiating odontogenic keratocyst (OKC) from other radiolucent jaw lesions like ameloblastoma is clinically important but radiographically difficult. Recent advances in artificial intelligence (AI) show promise for enhancing diagnosis using cone‐beam computed tomography (CBCT). This study aims to systematically evaluate and meta‐analyze the diagnostic accuracy of AI models in detecting OKCs on CBCT imaging. Methods A systematic review and meta‐analysis was conducted according to PRISMA‐DTA guidelines. Five electronic databases were searched through July 6, 2025. Studies employing AI models for OKC detection using CBCT were included. Methodological quality was assessed using QUADAS‐2. Pooled estimates were computed using a random‐effects model, with heterogeneity evaluated via I2 and meta‐regression. The Eager test and funnel plot were employed to assess publication bias. Results Twelve studies were included. AI models demonstrated high diagnostic accuracy, characterized by a pooled sensitivity of 89% (95% CI: 79%–95%) and specificity of 92% (95% CI: 81%–97%), both exceeding 85%, along with a substantial diagnostic odds ratio (87.06) and a robust discriminative ability (AUC = 0.828). Deep learning (DL) models achieved higher sensitivity (91%) than machine learning (ML) models (86%), while ML models showed slightly higher specificity. Heterogeneity was substantial (I2 = 78%–93%). Publication year explained 57.2% of the variability in sensitivity. Conclusions AI‐based models, particularly CNN‐based DL architectures, demonstrate clinically relevant diagnostic performance with high sensitivity, specificity, and diagnostic odds ratios, supporting their potential role as adjunctive tools in CBCT‐based differentiation of OKCs from other odontogenic lesions.
Vertical root fracture (VRF) remains a diagnostic challenge because radiographic signs are often subtle, non-specific and influenced by imaging modality, fracture characteristics, artefacts and observer interpretation. This systematic review and meta-analysis assessed the accuracy of artificial intelligence (AI)-based models for VRF detection on cone-beam computed tomography (CBCT), periapical radiographs and panoramic radiographs and evaluated the certainty of evidence. This review followed Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy Studies recommendations and was prospectively registered in International Prospective Register of Systematic Reviews. Studies evaluating AI-based models for VRF detection on dental imaging were included. Sensitivity and specificity were pooled using bivariate random-effects models when complete 2 × 2 contingency tables were available. Summary receiver operating characteristic curves were generated. Risk of bias was assessed using Quality Assessment of Diagnostic Accuracy Studies-2 and certainty of evidence was evaluated using the Grading of Recommendations Assessment, Development and Evaluation approach adapted for diagnostic test accuracy. Post-hoc sensitivity analyses retained one representative dataset per study to assess the influence of multiple model-specific datasets. Six studies were included, of which 4 contributed to quantitative synthesis. The CBCT-based AI models showed high diagnostic performance under predominantly controlled or partially controlled conditions, with a pooled sensitivity of 85.3% and specificity of 89.0% (area under the curve [AUC]=0.921). The CBCT convolutional neural network–only subgroup showed sensitivity of 88.3% and specificity of 87.6% (AUC=0.929). Periapical radiograph-based models showed high sensitivity but limited specificity, with pooled sensitivity of 86.9% and specificity of 61.2% (AUC=0.708). The probabilistic neural network-only periapical subgroup improved sensitivity, but specificity remained limited. Sensitivity analyses confirmed preserved CBCT performance and persistently low periapical specificity. Panoramic radiography was assessed qualitatively because complete 2 × 2 data were unavailable. Certainty of evidence was moderate for CBCT and low for periapical radiographs. AI-based models show promising diagnostic potential for VRF detection, particularly on CBCT. However, current evidence mainly reflects performance under controlled or partially controlled conditions and should not be interpreted as definitive real-world diagnostic effectiveness. AI should currently be considered an adjunctive decision-support tool rather than an autonomous diagnostic method, especially for 2-dimensional imaging modalities. Prospective, multicentre clinical validation studies are needed.
Karla-Nogueira Matos, Hugo Henrique dos Santos Dantas Guimarães, Pedro Vitor Dos Santos Sobrinho et al.· European Endodontic Journal· 0 citations
Abstract Background Osteonecrosis of the femoral head (ONFH) is a common cause of hip disability in clinical practice. Early and accurate diagnosis can delay or even halt disease progression. In recent years, AI models based on medical imaging have been increasingly applied to the diagnosis of ONFH; however, a systematic evaluation of their diagnostic accuracy remains lacking. Objective This study aims to synthesize the overall diagnostic accuracy of medical imaging-based AI models for ONFH and to inform clinical decision-making. Methods This systematic review was conducted in accordance with the PRISMA-DTA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy Studies) guidelines and was prospectively registered in PROSPERO (CRD420261307216). We searched PubMed, Embase, Cochrane Library, and Web of Science up to March 8, 2026. Studies developing or validating AI models for ONFH diagnosis using imaging data were eligible. Risk of bias was assessed using the QUADAS-2 tool. Sensitivity, specificity, positive likelihood ratio (PLR), negative likelihood ratio (NLR), and diagnostic odds ratio (DOR) were pooled using a bivariate mixed-effects model, and a summary receiver operating characteristic (SROC) curve was constructed. Subgroup analyses were stratified by imaging modality (x-ray vs MRI), disease stage (early-stage ONFH vs all-stage ONFH), diagnostic criteria (Association Research Circulation Osseous [ARCO] staging vs other criteria), control group type (healthy controls vs disease controls), validation method (internal validation vs external validation), center type (single-center vs multicenter), and model type (deep learning vs machine learning). Meta-regression was performed to quantify the contribution of each covariate to between-study heterogeneity. Sensitivity analysis and Deeks asymmetry test assessed the robustness of the results and publication bias. Clinical utility was evaluated using the Fagan nomogram. Results A total of 12 studies comprising 16,189 hip joints were included. The pooled sensitivity was 0.91 (95% CI 0.87‐0.95), the pooled specificity was 0.95 (95% CI 0.93‐0.96), and the SROC AUC was 0.97 (95% CI 0.95‐0.98). Substantial between-study heterogeneity was observed (I²=72%, 95% CI 38%‐100%). Subgroup analysis showed that MRI-based models yielded a higher diagnostic odds ratio (DOR; 382, 95% CI 220‐665) than x-ray-based models (106, 95% CI 60‐190), while models that underwent external validation had a lower DOR (129, 95% CI 51‐329) than those with only internal validation (230, 95% CI 104‐510). Meta-regression identified imaging modality as the primary source of heterogeneity, explaining 92.1% of the between-study variance. Conclusions AI models demonstrate high diagnostic accuracy in imaging-based ONFH diagnosis. However, the current evidence is constrained by the limited number of included studies, predominantly retrospective designs, and a lack of adequate external validation, and should therefore be interpreted with caution. Future research should adopt multicenter prospective designs, standardize reference standards, and implement rigorous external validation to facilitate clinical translation.
FeiLong Lu, Li-Rong Wang, Wen-Bin Zhang et al.· Journal of Medical Internet...· 0 citations
This systematic review evaluated externally validated, clinician-comparative studies of artificial intelligence (AI)-based fracture detection on conventional radiography and computed tomography (CT) and summarized diagnostic performance by imaging modality.
We conducted a systematic review of diagnostic accuracy studies in PubMed/MEDLINE, Embase, and Latin American and Caribbean Health Sciences Literature without language or date restrictions. Eligible studies evaluated AI tools applied directly to radiographs or CT images for fracture detection in living human participants. The focused synthesis included studies with external validation and formal human-reader comparison. Risk of bias and applicability were assessed using the Quality Assessment of Diagnostic Accuracy Studies 2 Tool. Sensitivity and specificity were summarized descriptively and, when complete 2 × 2 data were available, exploratory univariate random-effects meta-analyses were performed separately for radiography and CT. Certainty of the evidence was assessed using the Grading of Recommendations Assessment, Development and Evaluation approach (GRADE).
Nineteen studies (13 radiography and 6 CT studies) met the inclusion criteria. Anatomical targets included facial bones, upper and lower extremities, pelvis and hip, ribs, and vertebral fractures. Overall risk of bias was low in 2 studies, unclear in 11, and high in 6, with patient selection being the main concern. Four radiography studies contributed complete 2 × 2 data, yielding a pooled sensitivity of 0.86 (95% confidence interval [CI], 0.71– 0.94) and a specificity of 0.84 (95% CI, 0.77–0.89). Three CT studies contributed complete 2 × 2 data, yielding a pooled sensitivity of 0.93 (95% CI, 0.90–0.95) and a specificity of 0.92 (95% CI, 0.81–0.97). GRADE certainty was low for pooled sensitivity and specificity in both modalities and very low for AI-assisted interpretation versus unaided human readers.
Externally validated AI systems for fracture detection on radiography and CT showed high diagnostic performance and often performed comparably to human readers. However, certainty remains limited by heterogeneity, small numbers of studies with complete 2 × 2 data, and patient-selection concerns. Current evidence supports AI as an assistive tool, but prospective clinically integrated validation is needed before broad implementation.
Oscar Luis Castro Guerrero, Jesús Yesith Goenaga Fruto, Laurens Natan Antonio Vásquez et al.· Indian Journal of Musculoske...· 0 citations
OBJECTIVES
To evaluate diagnostic performance of artificial intelligence (AI) models in detecting proximal caries across different radiographic modalities.
METHODS
We systematically searched in five electronic databases: Web of Science, PubMed, IEEE Xplore, ScienceDirect and CNKI. Essential study characteristics, AI models and their accuracy metrics for proximal caries detection were extracted. Methodological quality of eligible studies was assessed with QUADAS-2. Studies judged to be of adequate quality were retained for meta-analysis.
RESULTS
Twenty studies were included: fifteen employed bitewing radiographs, three utilized panoramic radiographs, and two adopted periapical radiographs. AI accuracy ranged from 28.5% to 100%. Due to the unavailability of essential 2 × 2 contingency data, only ten studies (six bitewing, three panoramic, and one periapical) were eligible for meta-analysis, yielding a pooled sensitivity of 76% (95% CI: 70%-80%), specificity of 94% (95% CI: 90%-96%), and Summary Receiver Operating Characteristic (SROC) AUC of 0.90 (95% CI: 0.87-0.92). Substantial heterogeneity was observed.
CONCLUSION
AI models show promising diagnostic accuracy for proximal caries, with bitewing radiographs yielding slightly better performance than panoramic views. While AI holds potential as a clinical decision-support tool, high heterogeneity and limited external validation remain significant barriers to clinical translation. Future work should prioritize prospective, multicenter validation and open datasets to support clinical translation.
ADVANCES IN KNOWLEDGE
This systematic review covers studies up to October 2025, simultaneously evaluate AI performance of different modalities, enabling direct cross-modal comparison. By stratifying analyses by task type, dataset size, and imaging modality, this study identifies specific sources of performance heterogeneity that have previously confounded pooled estimates in the field. This review also establishes evidence thresholds for clinical implementation, serving as a methodological checkpoint to guide future research.
Yiyang Wang, Di Fu, Ge Zhou et al.· Dento maxillo facial radiolo...· 0 citations
OBJECTIVE
This study aimed to evaluate the structural characteristics of mandibular alveolar bone in patients with Type 1 diabetes mellitus (T1DM), Type 2 diabetes mellitus (T2DM), and systemically healthy controls using panoramic radiography-based radiomic analysis combined with machine learning algorithms.
MATERIALS AND METHODS
A total of 225 panoramic radiographs (75 T1DM, 75 T2DM, 75 healthy controls) were retrospectively analyzed. ROIs were segmented from eight anatomical mandibular segments per subject, and 107 radiomic features were extracted using PyRadiomics. Interobserver reliability was confirmed by two-way random-effects ICC (≥0.85). A leakage-free pipeline was applied. Four machine learning algorithms were evaluated: Random Forest, ExtraTrees, SVM-RBF, and Logistic Regression.
RESULTS
Significant differences were identified in age (Kruskal-Wallis p = 0.0003) and sex (χ²=17.857, p = 0.0001). The best single-segment performance was achieved in the left mandibular corpus with Logistic Regression (Accuracy=0.833, F1=0.832, AUC=0.958). All segments showed significant radiomic differences (FDR q < 0.001). Sensitivity of 1.000 was achieved for T1DM and AUC=1.000 for T2DM. The feature glszm_SizeZoneNonUniformityNormalized showed the strongest discriminative power (H = 86.928, ε²=0.578).
CONCLUSION
Panoramic radiography-based radiomic analysis demonstrates high diagnostic performance in non-invasively distinguishing mandibular bone alterations among T1DM, T2DM, and healthy individuals, with potential as a clinical bone monitoring tool.
INTRODUCTION
Accurate localization of the mandibular canal in Cone-Beam Computed Tomography (CBCT) images is critical for preventing iatrogenic nerve injury during maxillofacial surgery and dental implant procedures. This systematic review and meta-analysis aimed to evaluate the diagnostic performance, anatomical localization accuracy, and time efficiency of deep learning-based artificial intelligence (AI) systems in automated mandibular canal segmentation compared to traditional manual expert annotations.
MATERIALS AND METHODS
A comprehensive literature search was conducted across PubMed, Scopus, Web of Science, IEEE Xplore, and Embase databases in accordance with PRISMA guidelines. Studies evaluating the performance of AI models for mandibular canal detection on CBCT scans using expert annotations as the reference standard were included. The primary outcome measure was the Dice Similarity Coefficient (DSC), while secondary outcomes included Average Symmetric Surface Distance (ASSD) and processing time. Statistical analyses were performed using a random-effects model.
RESULTS
A total of 38 unique studies comprising over 8,420 CBCT volumes were included in the quantitative synthesis. The pooled DSC for AI-driven segmentation was calculated as 0.82 (95% CI: 0.79-0.85). Subgroup analyses revealed that transformer-based architectures (DSC: 0.89) demonstrated significantly superior performance compared to traditional convolutional neural networks (CNNs). The pooled ASSD exhibited a high anatomical accuracy of 0.42 mm (95% CI: 0.38-0.47), which is close to voxel dimensions. Furthermore, the autonomous segmentation process was completed in an average of 32 seconds, whereas manual expert annotation took 600 seconds (p < 0.001), confirming an 18.7-fold timesaving in the clinical workflow.
DISCUSSION
Deep learning algorithms provide highly accurate, reproducible, and time-efficient results at a human-expert level in the automated segmentation of the mandibular canal on CBCT images. The integration of these AI systems into clinical protocols has the potential to enhance surgical safety and standardize preoperative planning processes in dental implantology.
Ramazan Ağırağaç· Journal of Stomatology Oral...· 0 citations
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.