2026· IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing· Vol 19, pp. 27633-27646· 0 citations· 48 references
TL;DR
This study constructs the first land-oriented remote sensing SGG dataset by integrating and refining land-scene samples from ReCon1M and satellite-based terrain and relationship and proposes a semantic–visual collaborative SGG framework, which combines oriented object detection, global contextual modeling, and semanticprior fusion to alleviate semantic–visual conflicts in remote sensing scenes.
Abstract
High-resolution remote sensing image interpretation is evolving from object-level perception toward semantic cognition. However, existing scene graph generation (SGG) methods are difficult to adapt to remote sensing imagery due to the lack of dedicated land-oriented benchmarks, semantic–visual inconsistency, and severe long-tailed relationship distributions. To address these issues, this study constructs the first land-oriented remote sensing SGG dataset by integrating and refining land-scene samples from ReCon1M and satellite-based terrain and relationship, containing 19 658 images, 50 object categories, and 45 relationship categories. Furthermore, a semantic–visual collaborative SGG framework is proposed, which combines oriented object detection, global contextual modeling, and semanticprior fusion to alleviate semantic–visual conflicts in remote sensing scenes. In addition, a dual-level prototype-constrained relation learning strategy is introduced to improve rare relationship recognition under long-tailed distributions. Experimental results show that, compared with PE-Net, the proposed method improves $R@100$/$mR@100$ from 43.59/22.75 to 52.81/36.51 under scene graph classification and from 15.09/5.10 to 17.60/8.94 under scene graph detection. The proposed dataset and framework provide an effective benchmark and methodological paradigm for high-level semantic understanding of remote sensing imagery.
Remote sensing scene graph generation (RS-SGG) aims to advance remote sensing image interpretation from primitive entity recognition to high-level holistic scene understanding. Due to the large spatial coverage of remote sensing images, objects are often organized into multiple functional subscenes, while predicate sem...
Wen-Bin Wang, Yi-Heng Chen, Hang Sun et al.· IEEE Transactions on Geoscie...· 0 citations
Open-vocabulary semantic segmentation (OVSS) of remote sensing faces severe performance degradation when encountering unseen scene distributions caused by geographic, sensor, and resolution variations. Existing vision–language approaches provide strong semantic priors but lack scene-invariant structural representations...
Wu-Biao Huang, Hu-Chen Li, Shuai Zhang et al.· IEEE Transactions on Geoscie...· 0 citations
HDSMNet is proposed, a dual-branch multimodal semantic segmentation network designed for optical–nDSM data that enhances discriminative dense feature representations in high-resolution remote sensing images through interaction with a compact set of geometry-guided anchors.
Han-Xun Gu, Jiang-Jie Hu, Li Wang et al.· Remote Sensing· 0 citations
Due to the severe scale variation of targets in remote sensing images, the dense distribution of objects, and the fact that many small targets occupy only a very limited number of pixels, existing detection methods are prone to losing shallow details and suffering from insufficient low-level semantic representation dur...
Fa-Quan Song, Wu Le, Ming Lv et al.· IEEE Transactions on Geoscie...· 1 citation
Experiments show that DGSRef improves diverse segmentation architectures with limited additional computation and parameters, confirming its effectiveness as a lightweight decoupled refinement framework.