A semantic segmentation(SS) method for remote sensing(RS) images based on lightweight attention(LA) U-Net
Abstract
High-resolution remote sensing(HRRS) image semantic segmentation is a key technical link in fields such as urban precision management, national land space survey, ecological environment monitoring, and disaster emergency response. In response to the problems of redundant parameters, high computational complexity, and insufficient multi-scale feature capture ability of the classic U-Net model in remote sensing interpretation tasks, which make it difficult to be deployed in resource-constrained devices like unmanned aerial vehicles and embedded systems in real time, this paper proposes a lightweight attention U-Net remote sensing image semantic segmentation method for edge deployment. This method is based on the U-Net encoding-decoding architecture and replaces the standard convolution with depthwise separable convolution to construct lightweight encoding and decoding units, minimizing the model’s parameters and computational costs. A channel-spatial mixed attention module adapted to the lightweight architecture is designed and embedded in the bottleneck layer and skip connection structure of the model to enhance the focusing ability on key feature elements and the multi-scale feature fusion effect, alleviating the problems of large-scale differences in remote sensing image features and strong interference from complex backgrounds. At the same time, a combined loss function combining cross-entropy loss and Dice loss is constructed to improve the inherent class imbalance problem of remote sensing images and enhance the segmentation accuracy of small target features. The experimental results of the system conducted on the publicly available Vaihingen and Potsdam high-resolution remote sensing datasets show that the proposed model has only 2.82M parameters, 63.3% less than the original U-Net, and only 12.15G of floating-point operations. On the Vaihingen dataset, the average intersection-over-union (mIoU) reaches 80.24%, and on the Potsdam dataset, mIoU reaches 78.97%. The IoU for small targets, such as cars, is improved by 4.98 percentage points compared to the original U-Net. Additionally, the inference speed for a 512×512 pixel image is 1.9 times that of the original U-Net, achieving an optimal balance between model lightweighting and segmentation accuracy This method can well adapt to the real-time interpretation requirements of remote sensing images in resource-constrained scenarios such as unmanned aerial vehicle airborne terminals and embedded terminals, and can provide efficient and reliable technical support for remote sensing intelligent applications in related industries.