Comparative Analysis of AI Architectures for Prostate Gland Segmentation on MRI
Abstract
Prostate gland segmentation on magnetic resonance imaging (MRI) is important for prostate cancer treatment planning, but manual segmentation is time-consuming and subject to inter-reader variability. Deep learning models may offer automation and efficiency, but boundary uncertainty, tissue heterogeneity, morphological gland variations, and image scanner variations can affect performance. We studied the generalizability of three pretrained models: Prostate MRI Anatomy (PMRIA), TotalSegmentator, and MedSAM, using T2-weighted images from 39 patient studies in a publicly available dataset. Prostate MRI Anatomy and TotalSegmentator were evaluated on three-dimensional MRI volumes. MedSAM, however, was applied independently to each axial slice using a standardized bounding-box prompt, and the resulting masks were reconstructed into a three-dimensional segmentation volume. Dice coefficient (DSC), sensitivity, specificity, and Hausdorff distance (HD) were computed from the complete 3D T2- weighted MRI volumes. For PMRIA, DSC was 0.780, sensitivity 0.915, specificity 0.993, and HD 115.034 mm. For TotalSegmentator, DSC was 0.827, sensitivity 0.816, specificity 0.997, and HD 9.85 mm. For MedSAM DSC was 0.421, sensitivity 0.938, specificity 0.953, and HD 35.6 mm. Pairwise differences were statistically significant for all comparisons except sensitivity between MedSAM and PMRIA. TotalSegmentator demonstrated the strongest performance, while PMRIA displayed several boundary outliers and MedSAM had a high false positivity. Overall, the performance varied substantially among the models depending on their specific architectures and implementations. Future work using larger, independent datasets is needed to confirm these results. Evaluating clinical performance metrics and dosimetric assessments will help determine the model’s clinical utility.