Sep 2026· Signal, Image and Video Processing· Vol 20· 0 citations· 42 references
TL;DR
MCMB-UNet is proposed, an effective dual-path encoder model with multi-attention mechanisms that achieves a favorable accuracy-efficiency trade-off compared to mainstream Transformer-based models, and demonstrates good applicability on the Inria Aerial Labeling and Massachusetts Buildings datasets.
Accurate extraction of the spatial distribution of buildings from remote sensing imagery in complex urban environments is essential for urban planning and development. However, existing methods often suffer from high computational costs and insufficient building boundary recovery, making it difficult to achieve both ef...
Yao Lu, Gang Cheng, Guo-Sheng Cai et al.· Italian National Conference...· 0 citations
Road extraction from high-resolution remote sensing images is crucial for urban planning and geographic information systems (GIS). However, complex background interference, severe occlusions, and the inherent morphological complexity of roads often lead to discontinuities and insufficient accuracy in extraction results...
Jia-Jia Liu, Xuan Zhao, Wen-Xiang Dong et al.· Frontiers in Computing and I...· 0 citations
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
This work proposes a lightweight dual-domain attention aggregation network (LDANet), aiming to achieve image super-resolution with both high efficiency and high quality, and proposes the pixel-embedding channel attention module, which achieves cross-channel global context awareness by jointly modeling pixel-level spati...
Wei Xue, Meng-Cheng Ma, Bing-Wen Hu et al.· ACM Transactions on Multimed...· 0 citations
Referring remote sensing image segmentation (RRSIS) commonly treats language as a fixed query that only modulates visual features. This open-loop design is brittle in aerial scenes containing repeated objects, weak appearance cues, and relational expressions: visual evidence cannot revise which words and relations shou...
Chongyang Li, Chen Wang, Wen-Kai Zhang et al.· IEEE Geoscience and Remote S...· 0 citations
The gated multi-scale interaction network (GMSINet), a U-shaped encoder–decoder framework with a hierarchical shifted-window self-attention Transformer backbone with a hybrid loss function is introduced to balance pixel-level supervision stability and region-level structural consistency, is proposed.
Guobiao Yao, Ze-Yu Zhang, Qing-Dong Wang et al.· Photogrammetric Engineering...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.