Skip to content
Open access

Seeing With Words: Autocaption-Guided Graph Transformer for Remote Sensing Segmentation

2026 · IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing · Vol 19, pp. 30015-30031 · 0 citations · 47 references

Abstract

Semantic segmentation of remote sensing (RS) imagery is a cornerstone of geospatial analysis, which supports applications, such as land cover mapping, urban planning, and environmental monitoring. Despite significant progress with neural networks, existing approaches remain limited by their reliance on visual features alone, which often proves insufficient for distinguishing categories with similar textures or spectral signatures. In this work, we propose a novel RS segmentation framework that sees with words by integrating automatically generated image captions into a new form of graph-based reasoning. Unlike conventional graph methods that rely solely on spatial or visual similarity, we introduce a caption-aware graph construction strategy that incorporates high-level semantic cues to define patch connectivity. This novel reasoning mechanism allows the graph transformer to jointly capture spatial structure and language-guided semantics, enabling more context-aware feature aggregation. Caption features are reinjected during decoding via cross-attention to ensure that semantic guidance is preserved even when spatial similarity dominates neighbor selection. A lightweight decoder with refinement reduces patch artifacts and yields coherent high-resolution maps. Experiments on three benchmarks (LoveDA, Potsdam, and Vaihingen) show that the proposed framework achieves performance competitive with state-of-the-art methods while introducing a new language-guided paradigm for graph reasoning in RS image segmentation.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.