Skip to content

GPCNet: Grid-Parallel Compact Network With Kernel-Adaptive Feature Fusion for Semantic Segmentation of High-Resolution UAV Imagery

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 5637527-5637527 · 0 citations · 96 references

Abstract

In recent years, encoder–decoder architectures have driven remarkable progress in remote sensing semantic segmentation. However, achieving both high accuracy and computational efficiency for high-resolution UAV imagery remains challenging. For Vision Transformer (ViT)-based models, the quadratic computational and memory complexity of self-attention with respect to the number of tokens becomes prohibitive for large-scale UAV imagery. In addition, complex scenes with rich textures and fine-grained structures often exhibit significant intraclass appearance variations, which may degrade feature consistency and lead to object misclassification and boundary misalignment. To address these challenges, this article proposes GPCNet, an efficient semantic segmentation framework for high-resolution UAV imagery. First, we introduce foresight encoder sparsification (FES), which constructs a compact sparse encoder by evaluating parameter importance and removing redundant parameters prior to training. Second, we propose a grid-parallel encoding with stage-end semantic exchange (GPE-SSE) strategy to significantly reduce the computational cost of ViT-based encoders for high-resolution images while maintaining long-range semantic dependencies through cross-grid information exchange. Finally, a kernel-adaptive feature fusion (KAFF) module is designed for the semantic decoder, which generates spatially variable low-pass kernels (LPKs), high-pass kernels (HPKs), and dynamic dilated kernels (DDKs) to adaptively modulate high-frequency details and contextual receptive fields, thereby improving intraclass feature consistency and segmentation accuracy. The experimental results on three public high-resolution UAV image datasets demonstrate that our method maintains competitive segmentation accuracy while substantially improving inference efficiency. Specifically, with the compact MiT-B2 as the encoder, our GPCNet-P achieves 72.32% and 81.31% mIoU on the UAVid and VDD test sets, respectively, while reducing encoder parameters by 40% and the total FLOPs by 71.3% compared to the dense encoder. Our code is publicly available at https://github.com/Memoristor/GPCNet

View source

Similar papers

Sep 2026

ALSRFormer: An adaptive transformer with dynamic window attention and multi-scale deformable feed-forward network for remote sensing image segmentation.

High-resolution remote sensing images present considerable challenges for semantic segmentation due to their complex object structures and extensive spatial distribution. Effective segmentation requires capturing fine-grained local details while simultaneously modeling long-range dependencies. Convolutional Neural Netw...

Zi-Qi Jia, Jia-Wei Zhang, Dong-En Guo et al. · 0 citations
Sep 2026

MCMB-UNet: A dual-encoder network with multi-attention for remote sensing image segmentation

MCMB-UNet is proposed, an effective dual-path encoder model with multi-attention mechanisms that achieves a favorable accuracy-efficiency trade-off compared to mainstream Transformer-based models, and demonstrates good applicability on the Inria Aerial Labeling and Massachusetts Buildings datasets.

Dong-Dong Huang, Yu-Hong Ding · 0 citations
2026

AsyScale-Net: Fine-Grained Tiny Object Recognition Across Heterogeneous UAV–Satellite Imagery

Fine-grained classification of tiny objects in remote sensing (RS) is severely hindered by extreme ground sampling distance (GSD) variance and strict edge-computing constraints. To tackle these challenges, we propose AsyScale-Net. First, by leveraging comprehensive multifidelity imagery from our custom SkyView dataset,...

Rou Su, Pei-Jun Lee · 0 citations
Open access 2026

Dual-Level Prototype Alignment via Cross-Attention for Few-Shot Remote Sensing Semantic Segmentation

DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...

Mustafa Alawadi, M. Fateh · 0 citations
Open access Oct 2026

RL-Axial: Learning When to Compute Global Context for Multimodal Segmentation

High-resolution remote sensing semantic segmentation is a key enabler for fine-grained urban mapping and related geospatial applications. However, complex textures, severe scale variations, and cross-modal discrepancies make models with fixed architectures and computation paths struggle to combine reliable multimodal f...

Dong Xing, Hang Yang, Jin-He Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.