2026· IEEE Transactions on Geoscience and Remote Sensing· Vol 64, pp. 5637527-5637527· 0 citations· 96 references
Abstract
In recent years, encoder–decoder architectures have driven remarkable progress in remote sensing semantic segmentation. However, achieving both high accuracy and computational efficiency for high-resolution UAV imagery remains challenging. For Vision Transformer (ViT)-based models, the quadratic computational and memory complexity of self-attention with respect to the number of tokens becomes prohibitive for large-scale UAV imagery. In addition, complex scenes with rich textures and fine-grained structures often exhibit significant intraclass appearance variations, which may degrade feature consistency and lead to object misclassification and boundary misalignment. To address these challenges, this article proposes GPCNet, an efficient semantic segmentation framework for high-resolution UAV imagery. First, we introduce foresight encoder sparsification (FES), which constructs a compact sparse encoder by evaluating parameter importance and removing redundant parameters prior to training. Second, we propose a grid-parallel encoding with stage-end semantic exchange (GPE-SSE) strategy to significantly reduce the computational cost of ViT-based encoders for high-resolution images while maintaining long-range semantic dependencies through cross-grid information exchange. Finally, a kernel-adaptive feature fusion (KAFF) module is designed for the semantic decoder, which generates spatially variable low-pass kernels (LPKs), high-pass kernels (HPKs), and dynamic dilated kernels (DDKs) to adaptively modulate high-frequency details and contextual receptive fields, thereby improving intraclass feature consistency and segmentation accuracy. The experimental results on three public high-resolution UAV image datasets demonstrate that our method maintains competitive segmentation accuracy while substantially improving inference efficiency. Specifically, with the compact MiT-B2 as the encoder, our GPCNet-P achieves 72.32% and 81.31% mIoU on the UAVid and VDD test sets, respectively, while reducing encoder parameters by 40% and the total FLOPs by 71.3% compared to the dense encoder. Our code is publicly available at https://github.com/Memoristor/GPCNet
High-resolution remote sensing images present considerable challenges for semantic segmentation due to their complex object structures and extensive spatial distribution. Effective segmentation requires capturing fine-grained local details while simultaneously modeling long-range dependencies. Convolutional Neural Netw...
MCMB-UNet is proposed, an effective dual-path encoder model with multi-attention mechanisms that achieves a favorable accuracy-efficiency trade-off compared to mainstream Transformer-based models, and demonstrates good applicability on the Inria Aerial Labeling and Massachusetts Buildings datasets.
Dong-Dong Huang, Yu-Hong Ding· Signal, Image and Video Proc...· 0 citations
Fine-grained classification of tiny objects in remote sensing (RS) is severely hindered by extreme ground sampling distance (GSD) variance and strict edge-computing constraints. To tackle these challenges, we propose AsyScale-Net. First, by leveraging comprehensive multifidelity imagery from our custom SkyView dataset,...
Rou Su, Pei-Jun Lee· IEEE Geoscience and Remote S...· 0 citations
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
High-resolution remote sensing semantic segmentation is a key enabler for fine-grained urban mapping and related geospatial applications. However, complex textures, severe scale variations, and cross-modal discrepancies make models with fixed architectures and computation paths struggle to combine reliable multimodal f...
Dong Xing, Hang Yang, Jin-He Zhang et al.· Cognitive Computation· 0 citations