Skip to content
Open access

Prior-Enhanced Axial TransUNet: integrating MedSAM priors and sparse region-aware attention for kidney-tumor segmentation

Sep 2026 · Frontiers in Artificial Intelligence · 0 citations · 31 references

Abstract

Accurate kidney and renal-tumor segmentation is challenging because lesion size, location, morphology, and boundary contrast vary substantially across abdominal CT scans. Most existing methods rely on a single form of local evidence and struggle to recover the boundaries of small lesions while maintaining global anatomical consistency. This paper proposes Prior-Enhanced Axial TransUNet (PEAT), which combines the broad medical representations of foundation models with task-specific discriminative learning. PEAT introduces three core components: Bottleneck Semantic Alignment (BSA) uses the frozen MedSAM image encoder as a training-stage feature reference and guides the semantic organization of the task network's bottleneck representation through an L1 alignment constraint; Dual-path Attention-enhanced Fusion (DAF) adaptively fuses CNN local features with MaxViT global contextual features by cascading coordinate attention and SENetV2; and Bi-granular Region-Aware Attention (BRA) achieves efficient non-local context modeling through Top- k window routing and patch-level sparse interaction. The MedSAM branch participates only during training, and only the task-specific network runs at inference. Under a setting in which all compared methods adopt the same preprocessing and evaluation protocol, PEAT achieves kidney and tumor Dice scores of 97.69% and 91.83% on KiTS21 and 96.80% and 92.58% on KiTS23, respectively. Across all four tasks, PEAT attains the highest Dice among the compared methods; ablation studies further show that the progressive integration of BSA, DAF, and BRA yields consistent gains, with the complete model outperforming all single and pairwise configurations. These results indicate that combining training-stage broad feature guidance with task-specific representation learning effectively improves kidney and renal-tumor segmentation and provides a viable route for the coordinated design of foundation-model priors and lightweight task networks.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.