Skip to content
Open access

Knowledge Graph Enhanced for Zero-Shot Semantic Segmentation in Remote Sensing Imagery

Jul 2026 · ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences · Vol XI-3-2026, pp. 125-132 · 0 citations · 9 references

TL;DR

A Knowledge Graph (KG) enhanced ZSSS framework is proposed, which introduces explicit hierarchical and relational information into class embeddings to achieve more structured and semantically consistent representations.

Abstract

Abstract. Zero-shot semantic segmentation (ZSSS) is a crucial task in remote sensing image understanding, yet existing methods still suffer from limited generalization to unseen classes. To address this issue, we propose a Knowledge Graph (KG) enhanced ZSSS framework, which introduces explicit hierarchical and relational information into class embeddings to achieve more structured and semantically consistent representations. Specifically, a KG class encoder is designed, consisting of the class enhanced query (CEQ) and class enhanced embedding (CEE) modules, which extract class-relevant subgraphs from a self-constructing Remote Sensing Semantic Class Knowledge Graph (RSSCKG) and generate knowledge-enriched embeddings through a text encoder. Experiments on three public remote sensing datasets demonstrate that the proposed method consistently improves performance across seven state-of-the-art ZSSS frameworks. The integration of KG-based embeddings yields significant gains in the evaluation metrics, with particularly strong improvements on unseen classes, while maintaining accuracy on seen classes. Compared with enhancement strategies based on large language model (LLM) generated descriptions, the proposed KG class encoder exhibit superior semantic separability and stability. These results validate the effectiveness, generalization, and scalability of the proposed framework for ZSSS in remote sensing imagery.

Read PDF

Similar papers

Sep 2026

Multimodal graph-based fusion via image descriptions for few-shot open-set recognition.

A multimodal Graph-based Fusion (MGF) framework that learns visually grounded semantic representations to enhance FSOR performance and achieves superior open-set recognition and competitive closed-set classification performance.

Xilang Huang, Seon-Han Choi · 0 citations
2026

Graph2Scene: Generating Remote Sensing Imagery and Labels via Scene Graphs and Low-Rank Representation

High-precision, pixel-level annotations are indispensable for remote sensing semantic segmentation and related tasks, yet producing such labels manually is prohibitively expensive. Although recent generative models can synthesize realistic remote sensing data, existing approaches typically either rely heavily on preexi...

Shaoxuan Zhao, Xiao-Guang Zhou, Dong-Yang Hou et al. · 0 citations
2026

Enhancing Scene Generalization for Open-Vocabulary Remote Sensing Segmentation via Semantic–Structural Collaboration

Open-vocabulary semantic segmentation (OVSS) of remote sensing faces severe performance degradation when encountering unseen scene distributions caused by geographic, sensor, and resolution variations. Existing vision–language approaches provide strong semantic priors but lack scene-invariant structural representations...

Wu-Biao Huang, Hu-Chen Li, Shuai Zhang et al. · 0 citations
Open access Aug 2026

Soft Attention Enhanced CNN and LSTM Based Framework for Semantic Description Generation of Remote Sensing Imagery

Remote sensing images are complex, which makes it difficult to interpret and generate semantically appropriate textual description. To get a semantically relevant description, it is important to identify complex objects and understand the contextual relationships between them. In such cases, deriving contextually accur...

D. Pawade, Sonali Patil, R. Arya et al. · 0 citations
2026

Text-Guided Dual Refinement for Domain Generalized Semantic Segmentation in Remote Sensing

Recently, domain generalized remote sensing semantic segmentation (DG-RSSS) methods leverage vision foundation models (VFMs) with a parameter-efficient fine-tuning (PEFT) strategy to achieve remarkable progress. Although VFMs offer robust representations under distribution shifts between different remote sensing scenes...

Mu-Xin Liao, Mei-Ying Liao, Yu-Ting Sun et al. · 0 citations
Open access 2026

Dual-Level Prototype Alignment via Cross-Attention for Few-Shot Remote Sensing Semantic Segmentation

DLPANet is proposed, a novel dual-level prototype alignment network centered on Prototype-Guided Spatial Attention, enabling simultaneous modeling of scene context and fine-grained details and demonstrates that the decoupled dual cross-attention mechanism provides superior prototype-query alignment compared to prior gl...

Mustafa Alawadi, M. Fateh · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.