Skip to content
Review Open access

Artificial Intelligence for Diagnosis of Temporomandibular and Cranio-Cervico-Mandibular Musculoskeletal Disorders: A Systematic Review and Exploratory Diagnostic Test Accuracy Meta-Analysis

Aug 2026 · Diagnostics · Vol 16, pp. 2468 · 0 citations · 32 references
Medicine

TL;DR

Artificial intelligence-based methods demonstrate promising performance in selected image-based TMJ osteoarthritis tasks, but present evidence does not justify autonomous diagnosis or replacement of established clinical and imaging reference standards.

Abstract

Objectives: To systematically evaluate the diagnostic accuracy, clinical applicability, and methodological maturity of artificial intelligence (AI)-based methods for temporomandibular disorders (TMD), temporomandibular joint (TMJ) abnormalities, and related cranio-cervico-mandibular (CCM) musculoskeletal conditions compared with conventional diagnostic methods and accepted reference standards. Materials and Methods: This systematic review and exploratory diagnostic test accuracy meta-analysis was conducted in accordance with PRISMA 2020 and PRISMA-DTA. The protocol was retrospectively registered in PROSPERO (CRD420261428138). PubMed/MEDLINE, Embase, and Scopus were searched from database inception through February 2026. Eligibility for the primary synthesis was restricted to published studies in English or Spanish involving adults aged 18 years or older. All extracted records were re-audited article by article to align the evidence with the diagnostic question. The domain-specific quantitative synthesis was restricted to TMJ osteoarthritis studies with explicit 2 × 2 diagnostic data or a unique, verifiable reconstruction from reported class totals and sensitivity/specificity. Risk of bias was assessed with QUADAS-2. Results: From 1471 records identified, 174 entered the master extraction dataset. After reclassification, 84 records were retained for primary TMD/TMJ qualitative synthesis, 8 as secondary CCM musculoskeletal evidence, 31 as conventional or reference standard supporting evidence, 33 as methodological or contextual evidence, 4 as differential orofacial pain evidence, and 14 as excluded or minimal-background records. Twenty-one studies were assessed as potential diagnostic accuracy candidates. Three TMJ osteoarthritis studies contributed to the domain-specific exploratory meta-analysis: two with explicit 2 × 2 data and one with a reproducible reconstruction. Pooled sensitivity was 0.791 (95% CI: 0.700–0.861) and pooled specificity was 0.869 (95% CI: 0.811–0.911). Heterogeneity was substantial for sensitivity (I2 = 68.2%) and moderate for specificity (I2 = 57.3%). Conclusions: AI demonstrates promising performance in selected image-based TMJ osteoarthritis tasks. Nevertheless, the evidence remains exploratory because only three studies were quantitatively comparable, one table was reconstructed, and modalities and validation designs differed. AI should be interpreted as an augmentative decision support tool rather than a replacement for MRI, CBCT, or validated clinical frameworks such as DC/TMD. Clinical Relevance: AI may support image-based TMD/TMJ workflows, but present evidence does not justify autonomous diagnosis or replacement of established clinical and imaging reference standards.

Read PDF

Similar papers

Review Open access Aug 2026

Validity of Using Artificial Intelligence to Predict Skeletal Maturation in Orthodontic Treatment: A Systematic Review and Meta-analysis

Abstract This systematic review aimed to evaluate the performance and clinical validation of artificial intelligence (AI) and rule-based methods for skeletal maturation assessment using cervical vertebral maturation (CVM) and hand–wrist radiography. The review was conducted in accordance with the PRISMA 2020 guidelines and registered in PROSPERO. The PICOS framework guided selection of experimental, randomized controlled, and observational studies that applied AI to radiographic images of the cervical vertebrae or hand–wrist region. Eligible studies compared AI-derived classifications with expert assessments. Studies based on non-radiographic indicators, non-English publications, and investigations not involving AI methodologies were excluded. A systematic search was performed in PubMed, Scopus, EBSCOhost, and SpringerLink for articles published between 2015 and 2025. The methodological quality and risk of bias were evaluated using the QUADAS-2 tool. For CVM-based studies reporting sufficient quantitative data, a random-effects meta-analysis with restricted maximum likelihood (REML) estimation was performed using logit-transformed accuracy proportions. Heterogeneity was assessed using the Q statistic and the I 2 index. Studies based on hand–wrist radiography were excluded from the quantitative synthesis due to heterogeneity in outcome measures and the limited number of studies. Seven studies met the eligibility criteria, including five that used lateral cephalograms for CVM analysis and two that used hand–wrist radiographs for skeletal age estimation. CVM-based models demonstrated moderate to high diagnostic performance, with reported accuracy or agreement ranging from 60.4 to 82.8%, and kappa values reaching 0.985. These findings corresponded to a pooled accuracy of 80%. When compared with human observer assessments, several CVM-based models achieved comparable levels of agreement in selected studies. However, their performance remained variable across validation settings. Hand–wrist-based models showed consistently high agreement with human observers, with correlation coefficients up to r  = 0.98 and mean absolute errors below 6 months, indicating stable reproducibility relative to expert readings. AI demonstrates promising potential for assisting skeletal maturation assessment, although accuracy varies across model architectures and validation settings. Current evidence supports the use of AI primarily as a diagnostic support tool, with further prospective multicenter studies required before routine clinical implementation. However, the current evidence remains limited due to the small number of studies and methodological heterogeneity.

Gita Gayatri, Endah Mardiati, Ani Melani et al. · 0 citations
Review Open access Aug 2026

Diagnostic Accuracy of Medical Imaging–Based Artificial Intelligence for Osteonecrosis of the Femoral Head: Systematic Review and Meta-Analysis

Abstract Background Osteonecrosis of the femoral head (ONFH) is a common cause of hip disability in clinical practice. Early and accurate diagnosis can delay or even halt disease progression. In recent years, AI models based on medical imaging have been increasingly applied to the diagnosis of ONFH; however, a systematic evaluation of their diagnostic accuracy remains lacking. Objective This study aims to synthesize the overall diagnostic accuracy of medical imaging-based AI models for ONFH and to inform clinical decision-making. Methods This systematic review was conducted in accordance with the PRISMA-DTA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses of Diagnostic Test Accuracy Studies) guidelines and was prospectively registered in PROSPERO (CRD420261307216). We searched PubMed, Embase, Cochrane Library, and Web of Science up to March 8, 2026. Studies developing or validating AI models for ONFH diagnosis using imaging data were eligible. Risk of bias was assessed using the QUADAS-2 tool. Sensitivity, specificity, positive likelihood ratio (PLR), negative likelihood ratio (NLR), and diagnostic odds ratio (DOR) were pooled using a bivariate mixed-effects model, and a summary receiver operating characteristic (SROC) curve was constructed. Subgroup analyses were stratified by imaging modality (x-ray vs MRI), disease stage (early-stage ONFH vs all-stage ONFH), diagnostic criteria (Association Research Circulation Osseous [ARCO] staging vs other criteria), control group type (healthy controls vs disease controls), validation method (internal validation vs external validation), center type (single-center vs multicenter), and model type (deep learning vs machine learning). Meta-regression was performed to quantify the contribution of each covariate to between-study heterogeneity. Sensitivity analysis and Deeks asymmetry test assessed the robustness of the results and publication bias. Clinical utility was evaluated using the Fagan nomogram. Results A total of 12 studies comprising 16,189 hip joints were included. The pooled sensitivity was 0.91 (95% CI 0.87‐0.95), the pooled specificity was 0.95 (95% CI 0.93‐0.96), and the SROC AUC was 0.97 (95% CI 0.95‐0.98). Substantial between-study heterogeneity was observed (I²=72%, 95% CI 38%‐100%). Subgroup analysis showed that MRI-based models yielded a higher diagnostic odds ratio (DOR; 382, 95% CI 220‐665) than x-ray-based models (106, 95% CI 60‐190), while models that underwent external validation had a lower DOR (129, 95% CI 51‐329) than those with only internal validation (230, 95% CI 104‐510). Meta-regression identified imaging modality as the primary source of heterogeneity, explaining 92.1% of the between-study variance. Conclusions AI models demonstrate high diagnostic accuracy in imaging-based ONFH diagnosis. However, the current evidence is constrained by the limited number of included studies, predominantly retrospective designs, and a lack of adequate external validation, and should therefore be interpreted with caution. Future research should adopt multicenter prospective designs, standardize reference standards, and implement rigorous external validation to facilitate clinical translation.

FeiLong Lu, Li-Rong Wang, Wen-Bin Zhang et al. · 0 citations
Review Aug 2026

Morphological changes of temporomandibular joint architecture in patients with temporomandibular disorder: A systematic review and meta-analysis.

BACKGROUND Temporomandibular disorders (TMD) are frequently associated with structural alterations of the temporomandibular joint (TMJ). This systematic review evaluated morphological changes in TMJ components including condyle, articular disc, joint space, glenoid fossa, and articular eminence in individuals with TMD. METHODS The systematic review included studies done in adult subjects with any one sign or symptom of TMD or studies done in subjects diagnosed with TMD, not limited to DC/TMD criteria, RDC/TMD criteria, using three-dimensional imaging. Systematic searches were conducted in 8 databases from inception to December 2025. Analytical observational studies were selected, and critical appraisal was performed. Joanna Briggs Institute (JBI) guidelines for systematic effectiveness reviews were followed for data appraisal, extraction, and synthesis. RESULTS 190 articles are included in the systematic review. Most studies reported alterations in various components of the TMJ in patients with temporomandibular disorder. CONCLUSION TMD appears to be associated with significant osseous and soft-tissue changes, including condylar flattening, erosions, osteophyte formation, sclerosis, reduced condylar volume, glenoid fossa variations, altered articular eminence morphology and inclination, disc displacement, disc morphological alterations, and changes in joint space dimensions.

Sarika K, Vineetha Karuveetil, Ajithkumar V.V et al. · 0 citations
Review Open access Jul 2026

Artificial intelligence for assessment of the relationship between the mandibular canal and third molars using CBCT and panoramic radiography: a systematic review.

BACKGROUND Accurate assessment of the relationship between mandibular third molars and the mandibular canal is essential to reduce the risk of inferior alveolar nerve injury. Artificial intelligence (AI), including deep learning and convolutional neural network-based models, has been increasingly investigated using panoramic radiography and cone-beam computed tomography (CBCT) for automated diagnostic support in third molar assessment. METHODS A systematic review (SR) was conducted following the PECOS framework to answer the research question: "Can artificial intelligence determine the relationship between the mandibular canal and mandibular third molars?" Searches were performed in the PubMed, Scopus, CAPES Periodicals, LILACS, and Cochrane Library databases up to February 2026. Studies evaluating AI models for identifying or classifying the relationship between the mandibular canal and mandibular third molars were included. Due to methodological heterogeneity among studies, a qualitative synthesis was performed. After screening titles, abstracts, and full texts, 29 studies met the eligibility criteria, including 15 using panoramic radiographs, 9 using CBCT, and 5 evaluating both imaging modalities. Protocol registration was not performed. RESULTS AI-based models demonstrated favourable diagnostic performance for identifying and classifying the spatial relationship between the mandibular canal and mandibular third molars. Reported accuracy values frequently ranged from 80% to 97%, while sensitivity and specificity commonly exceeded 85%. Area under the ROC curve (AUC) values ranged from 0.84 to 0.98 across studies, indicating strong discriminatory capacity of CNN-based models for canal segmentation, spatial classification, and risk assessment tasks using panoramic radiographs and CBCT images. CONCLUSION Current evidence suggests that AI may support radiographic interpretation and pre-surgical assessment of the relationship between the mandibular canal and mandibular third molars. However, methodological heterogeneity and limited external validation highlight the need for further standardized multicenter studies before broader clinical implementation.

Melissa Gimenes Araújo, M. Miranda-Viana, L. M. P. Salzedas et al. · 0 citations
Review Open access Jul 2026

The diagnostic accuracy of artificial intelligence in detecting foot and ankle fractures: a systematic review and meta-analysis of diagnostic test accuracy.

INTRODUCTION Foot and ankle fractures, including radiographically subtle or occult injuries, present a diagnostic challenge in emergency settings, with missed diagnoses causing severe complications. Artificial intelligence (AI), specifically deep learning, offers a promising adjunct for radiographic interpretation. This systematic review and meta-analysis evaluates the diagnostic test accuracy of AI in detecting foot and ankle fractures and appraises the available evidence for occult fracture detection where reported. METHOD Adhering to PRISMA-DTA guidelines, a systematic search was conducted across PubMed, Embase, and Scopus from inception to 10 May 2026. Studies evaluating AI algorithms for identifying foot and ankle fractures on radiographs or CT were included, with occult-fracture data extracted separately when available. Extracted data populated 2 × 2 contingency tables. We utilized a bivariate random-effects model to calculate pooled sensitivity, specificity, and the summary receiver operating characteristic (SROC) curve. Quality was assessed via QUADAS-2. RESULTS Sixteen studies encompassing over 37,000 radiographs or images were included. Pooled sensitivity was 92.4% (95% CrI 82.5-96.1%) and pooled specificity was 95.1% (95% CrI 83.3-98.3%). The diagnostic odds ratio was 236.64 (95% CrI 33.56-1087.11). At a 25% pre-test probability, a positive AI result increased post-test probability to 86%, while a negative result reduced it to 3%. Substantial between-study heterogeneity was observed. Only two studies explicitly reported occult fracture cases; therefore, the pooled estimates primarily reflect AI performance for overall foot and ankle fracture detection rather than occult fractures alone. CONCLUSION AI algorithms demonstrate high diagnostic accuracy for foot and ankle fracture detection and show meaningful clinical utility as adjunctive tools. However, current evidence is insufficient to support robust occult-fracture-specific conclusions, and all included studies were retrospective. Prospective validation, standardized reporting of occult fracture subgroups, and real-world impact studies are needed before widespread clinical implementation.

Gregorius Thomas Prasetiyo, J. Wijaya · 0 citations
Aug 2026

Large language models demonstrate variable diagnostic performance and a systematic risk of undertriage in surgical triage of feline metacarpal and metatarsal fractures.

Objective To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures. Methods In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and metatarsal fractures were evaluated. Two board-certified veterinary orthopedic surgeons established a reference standard for surgical versus conservative management. Five LLMs (ChatGPT, version 5.2 [OpenAI Inc]; Gemini, version 3 Pro [Alphabet Inc]; Grok, version 4.1 [SpaceXAI]; Qwen, version 3.5 [Alibaba Cloud]; and Claude Sonnet, version 4.5 [Anthropic PBC]; and Claude Sonnet, version 4.5 [Anthropic PBC]) assessed anonymized case summaries using a standardized zero-shot prompt. Model recommendations were compared with the reference standard to calculate accuracy, sensitivity, specificity, and the Cohen κ. Results The reference standard classified 49 cases (67.1%) as surgical and 24 (32.9%) as conservative. ChatGPT achieved the highest performance (accuracy, 84.9%; sensitivity, 79.6%; specificity, 95.8%; κ = 0.69). Other models showed lower performance; Qwen and Claude Sonnet failed to identify any surgical cases (0% sensitivity). A systematic bias toward conservative management was observed across all models, causing high false-negative rates (undertriage). Conclusions LLMs demonstrate highly variable diagnostic performance and a systematic risk of undertriage in feline fracture assessment. While top-performing models approach expert-level agreement, others fail in critical clinical scenarios. Clinical Relevance Independent LLM use for surgical decision-making in feline orthopedics is not recommended due to undertriage risks. These tools require strict clinician oversight for preliminary triage.

S. Okur, Ç. Özkalıpçı, Büşra Baykal et al. · 0 citations