It is shown that conventional in-distribution evaluation can obscure clinically important generalization failures and support explicit cross-dataset testing as a key component of CHD segmentation evaluation, and conventional in-distribution evaluation can obscure clinically important generalization failures.
Abstract
Congenital heart disease (CHD) diagnosis and surgical planning often require patient-specific 3D anatomical models, but manual segmentation is labor-intensive, particularly in complex anatomies. Although deep-learning methods can automate this process, they are typically evaluated in-distribution, despite clinically relevant shifts in scanner, protocol, institution, population, and imaging modality. We present, to our knowledge, the first systematic evaluation of out-of-distribution (OOD) generalization in CHD segmentation, using ImageCHD as a held-out target cohort. We compare representative segmentation architectures under combined CT and CMR training, CT-only training, self-supervised pretraining, and limited target-domain adaptation. In-distribution performance proves to be a poor indicator of cross-cohort robustness: nnU-Net achieves the highest validation Dice (0.77) but falls to 0.51 on ImageCHD, while SwinUNETR generalizes substantially better, reaching 0.67 Dice. MAE and JEPA pretraining provide only modest additional benefit, suggesting that architecture contributes more to robustness than the tested pretraining strategies in this setting. When limited target-domain supervision is introduced, all SwinUNETR variants exceed 0.76 Dice with only 11 labeled ImageCHD cases. These findings demonstrate that conventional in-distribution evaluation can obscure clinically important generalization failures and support explicit cross-dataset testing as a key component of CHD segmentation evaluation.
A modality-routed 3D cardiac segmentation pipeline that combines TotalSegmentator-initialized nnU-Netv2 models with site-characterized, label-preserving appearance augmentation is proposed, suggesting that site-motivated appearance augmentation is a practical strategy for improving cross-site robustness in limited-data...
Tanishqua H Mudaliar, Justin D. Li, Daniel Lin et al.· 0 citations
It is suggested that ten annotated cases are sufficient for clinically useful segmentation, effectively reducing bottlenecks for both image annotation and training time.
S. Nagaraju, B. Abrahamsen, Ashkan Moradi et al.· 0 citations
MRD-UNet provides a practical balance between segmentation accuracy and computational efficiency and outperforms baseline CNNs and performs comparably to heavier transformer-based models while using significantly fewer parameters.
Musa Doğan, I. Ozkan· BMC Medical Imaging· 0 citations
As artificial intelligence becomes increasingly integrated into medical imaging practice, its robustness across heterogeneous real-world settings remains a major challenge. We quantified the effect of real-world distribution shifts on three-dimensional AI models for lung nodule analysis on CT and, motivated by these sh...
B. Bercean, Rafael Medelean, A. Tenescu et al.· Journal of imaging informati...· 0 citations
Background/Objectives: To develop and evaluate an automated CT-based framework for the quantitative assessment of fibrotic interstitial lung disease (ILD), including idiopathic pulmonary fibrosis (IPF), using a standardised six-level anatomical protocol and deep-learning lung segmentation. Methods: The segmentation dat...
Purpose Deep Learning (DL) has transformed cardiac image segmentation, yet its application to congenital heart disease (CHD) remains underexplored, with no prior systematic review in this emerging domain. This review evaluates current DL approaches for CHD CT segmentation, identifies best performing model architectures...
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.