Jul 2026· Journal of Applied Remote Sensing· Vol 20, pp. 034506 - 034506· 0 citations· 56 references
Engineering
TL;DR
Experimental results demonstrate that StructPointNet delivers competitive segmentation accuracy, and the model proves to be scalable and hardware-friendly, offering a practical paradigm for efficient multimodal interpretation in resource-limited settings.
Abstract
Abstract. Semantic segmentation is widely regarded as one of the most advanced scene understanding techniques in remote sensing. Although monomodal segmentation models have advanced recently, the rich potential of multimodal data is still largely untapped. Current multimodal approaches are often rigid, typically limited to bimodal inputs, and failing to adapt flexibly when three or more co-registered two-dimensional grid-based modalities are available. A more pressing issue is the computational burden inherent in traditional fusion paradigms; as the number of input modalities increases, parameter counts and processing loads often increase substantially, making them less practical for scalable multi-source remote sensing interpretation. Thus, creating a scalable and lightweight framework for co-registered grid-based multimodal remote sensing data remains an open challenge. To bridge this gap, we introduce StructPointNet, a lightweight segmentation framework built on the philosophy of “explicit extraction and reuse of geometric priors.” Our approach balances efficiency with precision through three core mechanisms: the structure-sensitive modality encoder, which captures modality-specific high-frequency geometric details via parallel Sobel branches and edge-guided attention; the Heterogeneity rectification layer and pyramidal hybrid backbone, which map diverse features into a shared latent space for global-local context modeling; and the boundary-aware point decoder, which refines boundary segmentation by resampling shallow structural features based on uncertainty estimates. Thanks to these designs, StructPointNet functions as a flexible framework for monomodal and multimodal segmentation with variable numbers of co-registered grid-based inputs while keeping the additional cost of each lightweight modality-specific branch controlled compared with conventional multistream backbone replication. We benchmarked our approach against multiple representative state-of-the-art models from the last five years on two public multimodal datasets. Experimental results demonstrate that StructPointNet delivers competitive segmentation accuracy, particularly in defining complex geo-object boundaries. Moreover, the model proves to be scalable and hardware-friendly, offering a practical paradigm for efficient multimodal interpretation in resource-limited settings.
Experiments show that DGSRef improves diverse segmentation architectures with limited additional computation and parameters, confirming its effectiveness as a lightweight decoupled refinement framework.
A Mahalanobis-Angle Boundary Loss (MABL) is proposed that explicitly enhances boundary and shape consistency and is introduced, built upon MABL, a boundary- aware remote sensing segmentation framework with Struc- tural Penalties.
Yue-Xi Song, Kai-Lai Sun, Zhuoyue Wang et al.· 0 citations
Despite the economic advantages of weakly supervised semantic segmentation (WSSS) in remote sensing imagery (RSI), existing VLM- and VFM-based approaches still struggle with domain-specific semantic ambiguities and geometric discontinuities. To address these challenges, we propose GeoSeC, the Geometric-Semantic Collabo...
Xiang-Rong Zhang, Jian-Xun Lai, Guan-Chun Wang et al.· IEEE Transactions on Image P...· 0 citations
The existing remote sensing image segmentation methods rely on predefined category labels, which are often insufficient to capture the complex spatial semantics inherent in geospatial concepts, such as flood inundation zones, landslide bodies, and industrial complexes. This letter presents ConceptSeg, a text-guided mul...
Yang Zhao, Ya-Wei Bai, Ming-Ming Jia et al.· IEEE Geoscience and Remote S...· 0 citations
Semantic segmentation is a fundamental task in remote sensing image (RSI) interpretation, aiming at pixel-wise classification of land cover. Recently, multimodal RSI semantic segmentation has attracted significant attention for its ability to alleviate the information bottleneck inherent in unimodal methods. However, e...
Guang-Yi Wei, Qian-Peng Chong, Zhen-Xi Wang et al.· IEEE Transactions on Geoscie...· 0 citations
DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...
Mustafa Alawadi, M. Fateh· Jordanian Journal of Compute...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.