SA-DeepLab: Enhancing DeepLabV3+ with Pixel-Wise Switchable Atrous Convolution for Adaptive Multi-Scale Semantic Segmentation
Abstract
Semantic segmentation assigns a semantic label to every pixel in an image, which is a fundamental task in computer vision with applications such as autonomous driving and medical imaging. Existing segmentation methods, whether CNN-based or transformer-based, both have limitations: the former are constrained by fixed dilation rates that hinder scale adaptation, while the latter suffer from quadratic computational complexity that prevents efficient deployment. To address these issues, this paper proposes SA-DeepLab, a CNN framework that combines a switchable atrous spatial pyramid pooling (SA-ASPP) module and a channel shuffle operation in the decoder. SA-ASPP employs a learnable spatial switch map to dynamically fuse two atrous convolutions with different dilation rates, enabling adaptive receptive field selection at each pixel, while a depthwise separable atrous convolution replaces the most dilated branch to reduce computational overhead. The channel shuffle operation rearranges feature channels across groups to encourage cross-group information exchange without extra parameters. Experiments on Pascal VOC 2012 and Cityscapes demonstrate that SA-DeepLab achieves competitive mIoU (86.30% and 81.78%) with a lightweight framework, outperforming both CNN-based and transformer-based competitors in the accuracy-efficiency tradeoff.