MGT–UNet: a hybrid CNN–transformer network with multi-scale feature learning and global context modeling for breast ultrasound segmentation
Accurate breast ultrasound image segmentation remains challenging because speckle noise, low contrast, and heterogeneous lesion appearance often degrade lesion boundary delineation. Convolutional neural networks provide effective local feature extraction but have limited capability for modeling long-range contextual information. Although transformer-based architectures improve global context modeling, effectively combining local multi-scale representations with global contextual information remains challenging. We propose MGT–UNet, a hybrid CNN–Transformer segmentation network that integrates multi-scale feature learning with global context modeling for breast ultrasound image segmentation. The encoder employs a Tri-Scale Context Extractor to learn hierarchical multi-scale representations, while a Global Context Transformer models long-range contextual dependencies at the bottleneck. The network was evaluated primarily on the BUSI breast ultrasound dataset using Dice coefficient, Intersection-over-Union, Hausdorff Distance at the 95th percentile, Precision, and Recall. Supplementary evaluations were conducted on the ISIC2018 dermoscopic dataset and a binary variant of the Synapse computed tomography dataset using the same segmentation protocol. To ensure reliability, performance metrics were averaged over multiple independent training runs with different random seeds. On the BUSI dataset, MGT–UNet achieved the highest Dice coefficient of 77.74%, outperforming representative CNN-, Transformer-, and hybrid segmentation models while also achieving the lowest HD $$_{95}$$ . The results indicate improved segmentation performance on the BUSI dataset, which contains images affected by speckle noise and weak boundary contrast. Evaluations on dermoscopic and computed tomography datasets complement these findings by examining the behavior of the proposed architecture across diverse modalities. Our results suggest that combining multi-scale feature learning with transformer-based global context modeling is beneficial for breast ultrasound segmentation. The proposed architecture provides a basis for further investigation of hybrid CNN–Transformer models in medical image segmentation.