Swin-DeepLabV3 is proposed, a hybrid semantic segmentation framework that integrates global and local feature modeling by combining a hierarchical Swin Transformer encoder with an Atrous Spatial Pyramid Pooling-based context module, demonstrating an effective balance between contextual representation and spatial precision for medical image segmentation tasks.
Abstract
Accurate medical image segmentation relies on precise boundary delineation and effective contextual modeling across multiple spatial scales. Conventional convolutional neural network-based architectures, such as U-Net, are effective at capturing local spatial details but are limited in modeling long-range dependencies due to their restricted receptive fields. Transformer-based approaches have recently shown potential in addressing this limitation by enabling global feature interactions; however, their direct application to dense prediction tasks often introduces high computational cost and challenges in multi-scale representation learning. This paper proposes Swin-DeepLabV3, a hybrid semantic segmentation framework that integrates global and local feature modeling by combining a hierarchical Swin Transformer encoder with an Atrous Spatial Pyramid Pooling-based context module. The Swin Transformer encoder employs shifted window self-attention to capture long-range dependencies while preserving hierarchical multi-scale representations. To enhance contextual understanding, the Atrous Spatial Pyramid Pooling module is incorporated at the bottleneck to aggregate spatial information at multiple dilation rates. A lightweight decoder with skip connections is then used to fuse high-level semantic features with low-level spatial details, supporting improved boundary localization. The proposed approach is evaluated on three public breast ultrasound datasets, namely BUSI, BUS-B, and BUS-BRA, under a 5-fold cross-validation protocol. Experimental results show that Swin-DeepLabV3 achieves strong overlap performance across datasets, with the best Dice scores of 83.54%, 88.49%, and 77.13% on BUS-B, BUS-BRA, and BUSI, respectively. The proposed architecture demonstrates an effective balance between contextual representation and spatial precision for medical image segmentation tasks. The implementation of the proposed method is publicly available at https://github.com/CaoMinhh/Swin-DeepLabV3
A tri-stream interaction paradigm replacing symmetric skip connections with directional fusion among semantic, spatial, and decoder-propagated streams at each decoding stage, and a Global Prototype Bank that captures dataset-level anatomical regularities via attention-based retrieval and gated EMA updates, providing pe...
Mohammed A. M. Elhassan, Qian-Fa Yuan, Zhizhong Xu et al.· Journal of King Saud Univers...· 0 citations
High-resolution remote sensing images present considerable challenges for semantic segmentation due to their complex object structures and extensive spatial distribution. Effective segmentation requires capturing fine-grained local details while simultaneously modeling long-range dependencies. Convolutional Neural Netw...
GLNet adopts a dual-branch encoder that combines a CNN-based Local Detail Perception Branch with a Mamba-based Global Context Modeling Branch, enabling the joint extraction of fine-grained local features and long-range semantic representations.
Medical image segmentation requires high accuracy and robustness, yet practical commercial deployment also demands privacy preservation and computational efficiency. In this context, the U-Net architecture, which can be inherently decoupled into independent encoder and decoder components, serves as a natural commercial...
Semantic segmentation assigns a semantic label to every pixel in an image, which is a fundamental task in computer vision with applications such as autonomous driving and medical imaging. Existing segmentation methods, whether CNN-based or transformer-based, both have limitations: the former are constrained by fixed di...
Objectives. We propose a modern method for semantic segmentation of ultra-high-resolution (4K) video frames in the field of Earth remote sensing using a modification of the DeepLabV3+ convolutional neural network.
Methods. The method is aimed at solving two critical problems: the limited receptive field of the model w...
A. A. Kozlov, S. Ablameyko· Informatics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.