Skip to content

StructPointNet: explicit geometric prior extraction and reuse for efficient multimodal remote sensing semantic segmentation

Jul 2026 · Journal of Applied Remote Sensing · Vol 20, pp. 034506 - 034506 · 0 citations · 56 references
Engineering

TL;DR

Experimental results demonstrate that StructPointNet delivers competitive segmentation accuracy, and the model proves to be scalable and hardware-friendly, offering a practical paradigm for efficient multimodal interpretation in resource-limited settings.

Abstract

Abstract. Semantic segmentation is widely regarded as one of the most advanced scene understanding techniques in remote sensing. Although monomodal segmentation models have advanced recently, the rich potential of multimodal data is still largely untapped. Current multimodal approaches are often rigid, typically limited to bimodal inputs, and failing to adapt flexibly when three or more co-registered two-dimensional grid-based modalities are available. A more pressing issue is the computational burden inherent in traditional fusion paradigms; as the number of input modalities increases, parameter counts and processing loads often increase substantially, making them less practical for scalable multi-source remote sensing interpretation. Thus, creating a scalable and lightweight framework for co-registered grid-based multimodal remote sensing data remains an open challenge. To bridge this gap, we introduce StructPointNet, a lightweight segmentation framework built on the philosophy of “explicit extraction and reuse of geometric priors.” Our approach balances efficiency with precision through three core mechanisms: the structure-sensitive modality encoder, which captures modality-specific high-frequency geometric details via parallel Sobel branches and edge-guided attention; the Heterogeneity rectification layer and pyramidal hybrid backbone, which map diverse features into a shared latent space for global-local context modeling; and the boundary-aware point decoder, which refines boundary segmentation by resampling shallow structural features based on uncertainty estimates. Thanks to these designs, StructPointNet functions as a flexible framework for monomodal and multimodal segmentation with variable numbers of co-registered grid-based inputs while keeping the additional cost of each lightweight modality-specific branch controlled compared with conventional multistream backbone replication. We benchmarked our approach against multiple representative state-of-the-art models from the last five years on two public multimodal datasets. Experimental results demonstrate that StructPointNet delivers competitive segmentation accuracy, particularly in defining complex geo-object boundaries. Moreover, the model proves to be scalable and hardware-friendly, offering a practical paradigm for efficient multimodal interpretation in resource-limited settings.

View source

Similar papers

Preprint Aug 2026

BASeg: Boundary-Aware Remote Sensing Segmentation with Structural Penalties

A Mahalanobis-Angle Boundary Loss (MABL) is proposed that explicitly enhances boundary and shape consistency and is introduced, built upon MABL, a boundary- aware remote sensing segmentation framework with Struc- tural Penalties.

Yue-Xi Song, Kai-Lai Sun, Zhuoyue Wang et al. · 0 citations
Sep 2026

GeoSeC: Geometric-Semantic Collaborative Learning for Weakly Supervised Remote Sensing Image Segmentation.

Despite the economic advantages of weakly supervised semantic segmentation (WSSS) in remote sensing imagery (RSI), existing VLM- and VFM-based approaches still struggle with domain-specific semantic ambiguities and geometric discontinuities. To address these challenges, we propose GeoSeC, the Geometric-Semantic Collabo...

Xiang-Rong Zhang, Jian-Xun Lai, Guan-Chun Wang et al. · 0 citations
2026

ConceptSeg: Zero-Shot Geospatial Concept Segmentation for Remote Sensing Images

The existing remote sensing image segmentation methods rely on predefined category labels, which are often insufficient to capture the complex spatial semantics inherent in geospatial concepts, such as flood inundation zones, landslide bodies, and industrial complexes. This letter presents ConceptSeg, a text-guided mul...

Yang Zhao, Ya-Wei Bai, Ming-Ming Jia et al. · 0 citations
2026

From Relation to Structure: Spatial-Semantic Guidance and Structure Refinement Network for Multimodal Remote Sensing Segmentation

Semantic segmentation is a fundamental task in remote sensing image (RSI) interpretation, aiming at pixel-wise classification of land cover. Recently, multimodal RSI semantic segmentation has attracted significant attention for its ability to alleviate the information bottleneck inherent in unimodal methods. However, e...

Guang-Yi Wei, Qian-Peng Chong, Zhen-Xi Wang et al. · 0 citations
Open access 2026

Dual-Level Prototype Alignment via Cross-Attention for Few-Shot Remote Sensing Semantic Segmentation

DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...

Mustafa Alawadi, M. Fateh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.