Deep learning-based segmentation of abdominal aortic aneurysm: a systematic review and meta-analysis of performance and clinical applicability
Abstract
Abdominal aortic aneurysm (AAA) management relies heavily on imaging for surveillance, treatment planning, and follow-up. Deep learning (DL)-based segmentation may improve the efficiency and reproducibility of AAA image analysis; however, reported performance varies across studies. This study aimed to systematically review and quantitatively summarize the performance of DL models for AAA-related segmentation. PubMed, Embase, Scopus, and Web of Science were searched for studies applying DL to AAA imaging in adults and reporting Dice score or intersection-over-union (IoU) as segmentation metrics. Each DL model evaluated on a held-out test set was treated as a separate model-level observation. Random-effects meta-analyses were used to pool Dice score and IoU. Subgroup analyses were performed by imaging modality, model category, and input dimensionality, and additional analyses provided structure-specific estimates for aneurysm sac, intraluminal thrombus (ILT), lumen, and EVAR-related segmentations. Across all targets, overall segmentation performance was 0.81 (95% CI 0.78–0.83) for Dice score and 0.72 (95% CI 0.67–0.77) for IoU. In subgroup analyses by model category, pooled Dice scores were 0.83 (95% CI 0.80–0.86) for U-Net family models, 0.82 (95% CI 0.77–0.86) for transformer or hybrid models, and 0.76 (95% CI 0.72–0.81) for other convolutional neural networks. By imaging modality, Dice scores were 0.87 (95% CI 0.84–0.89) for contrast-enhanced CT (CECT), 0.83 (95% CI 0.78–0.88) for non-contrast CT (NCCT), 0.84 (95% CI 0.81–0.87) for MRI, 0.78 (95% CI 0.74–0.82) for CT angiography (CTA), and 0.95 (95% CI 0.92–0.98) for X-ray. IoU subgroup analyses showed broadly similar patterns by modality and model category. In structure-specific analyses, aneurysm sac segmentation achieved a pooled Dice of 0.87 (95% CI 0.85–0.89) and IoU of 0.80 (95% CI 0.68–0.93); ILT segmentation reached 0.72 (95% CI 0.62–0.81) and 0.67 (95% CI 0.57–0.77); lumen segmentation 0.84 (95% CI 0.73–0.95) and 0.92 (95% CI 0.80–1.00); and EVAR-related targets 0.76 (95% CI 0.67–0.86) and 0.66 (95% CI 0.55–0.77) for Dice and IoU, respectively. DL-based models demonstrate good segmentation performance for the AAA sac and lumen, indicating clinically acceptable accuracy for applications such as volumetric surveillance and surgical planning. Performance for ILT and EVAR-related structures is more variable, suggesting that human oversight remains advisable for these targets in clinical practice. These findings support the use of DL-based segmentation in research and highlight its potential for selected clinical applications. However, the evidence base remains heterogeneous and is derived from a limited number of studies, with most models lacking external validation. Larger, standardized, multicentre evaluations with robust external validation are needed to confirm and refine these estimates.