ALSRFormer: An adaptive transformer with dynamic window attention and multi-scale deformable feed-forward network for remote sensing image segmentation.
Abstract
High-resolution remote sensing images present considerable challenges for semantic segmentation due to their complex object structures and extensive spatial distribution. Effective segmentation requires capturing fine-grained local details while simultaneously modeling long-range dependencies. Convolutional Neural Networks (CNNs) are well-suited for local detail extraction, whereas Vision Transformers (ViTs) excel at global dependency modeling, yet both exhibit inherent limitations. Although CNN-Transformer hybrid approaches have improved global feature representation, their reliance on fixed window partitioning, static feature fusion, and rigid convolutional kernels constrains performance on irregular objects and complex textures. To address these limitations, we propose the Adaptive Long Short Range Transformer (ALSRFormer) built upon the ConvNeXt backbone, forming the ConvALSR-Net model. Specifically, we design a Dynamic Window Multi-head Self-Attention (DynamicWMSA) mechanism that adaptively adjusts local attention windows according to image content. We further introduce a Multi-Scale Deformable Directional FFN (MSDD-FFN) to enhance the representation of complex textures and boundaries via multi-scale deformable convolutions. Additionally, Inception Depthwise Convolution (IDConv) and an AdaptiveFusion module are integrated to refine shallow features and improve global information fusion. Experiments on the LoveDA, Vaihingen, and Potsdam datasets show that ConvALSR-Net consistently outperforms existing methods, with mIoU improvements ranging from 0.6 to 1.1 percentage points over competitive baselines while using fewer parameters (65.15M) and lower FLOPs (70.30G) than the original ConvLSR-Net. The gains are more evident on the LoveDA dataset, where complex multi-scale scenes amplify the benefit of content-adaptive modeling.