HDSMNet is proposed, a dual-branch multimodal semantic segmentation network designed for optical–nDSM data that enhances discriminative dense feature representations in high-resolution remote sensing images through interaction with a compact set of geometry-guided anchors.
Abstract
High-resolution remote sensing semantic segmentation requires the joint modeling of local details, global semantics, and height-derived geometric structures, and it provides an important basis for urban object mapping, land-cover analysis, and fine-grained spatial understanding. However, in complex urban scenes, fine-grained boundaries, small objects, inter-class similarity, and spectral confusion can still weaken the stability of pixel-level prediction. To enhance discriminative dense feature representations in high-resolution remote sensing images, we propose HDSMNet, a dual-branch multimodal semantic segmentation network designed for optical–nDSM data. The network separately extracts appearance and semantic features from optical imagery and height–structural features from nDSM, and introduces a Height-Guided Sparse Cross-Modal Fusion (HGSCF) module. Rather than treating nDSM as an additional feature source for generic fusion, HGSCF derives contextual representations, local feature contrasts, and structural-discontinuity cues from encoded nDSM features and uses them to guide sparse anchor-based interaction between optical and height features. This design enhances discriminative dense feature representations through interaction with a compact set of geometry-guided anchors. To complement HGSCF at the output stage, HDSMNet further adapts a Context-Guided Refinement (CGR) path that combines intermediate-response-guided contextual aggregation with dynamic feature modulation. This supplementary path recalibrates decoder features for output refinement. Experiments on the ISPRS Potsdam and Vaihingen datasets show that HDSMNet achieves mIoU values of 86.57% and 84.22%, respectively; ablation results further identify HGSCF as the main contributor to the observed improvement.
Experiments show that DGSRef improves diverse segmentation architectures with limited additional computation and parameters, confirming its effectiveness as a lightweight decoupled refinement framework.
Optical–elevation data fusion is widely used in aerial remote sensing semantic segmentation, as optical imagery provides rich spectral and textural information, while DSM or DEM data offer complementary elevation-related structural cues. However, effective fusion remains challenging because optical and elevation repres...
Yi-Fan Yu, Song Deng, Yang Yang et al.· Remote Sensing· 0 citations
Semantic segmentation of high-resolution remote sensing images faces three major challenges in frequency-spatial feature fusion: background clutter mixed into high-frequency components, semantic discontinuities within large homogeneous regions, and loss of fine rigid boundaries caused by convolutional downsampling. Tra...
Qi-Yuan Zhang, Jian-Shun Liu· Italian National Conference...· 0 citations
Accurate landslide mapping from high-resolution remote sensing imagery requires both broad spatial context and precise boundary detail. State space models (SSMs), such as vision state space duality (VSSD), provide global receptive fields with linear complexity, but their global interaction mechanism offers no dedicated...
Accurate dam detection in high-resolution remote sensing imagery is challenging because dams exhibit substantial variations in scale and morphology, indistinct boundaries, and high visual similarity to roads, bridges, and shorelines. Existing single-modal detection methods rely primarily on spectral and texture informa...
Jian-Wei Zhu, De-Er Song, Tao-Hong Cai et al.· Remote Sensing· 0 citations
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.