Seeing With Words: Autocaption-Guided Graph Transformer for Remote Sensing Segmentation
Abstract
Semantic segmentation of remote sensing (RS) imagery is a cornerstone of geospatial analysis, which supports applications, such as land cover mapping, urban planning, and environmental monitoring. Despite significant progress with neural networks, existing approaches remain limited by their reliance on visual features alone, which often proves insufficient for distinguishing categories with similar textures or spectral signatures. In this work, we propose a novel RS segmentation framework that sees with words by integrating automatically generated image captions into a new form of graph-based reasoning. Unlike conventional graph methods that rely solely on spatial or visual similarity, we introduce a caption-aware graph construction strategy that incorporates high-level semantic cues to define patch connectivity. This novel reasoning mechanism allows the graph transformer to jointly capture spatial structure and language-guided semantics, enabling more context-aware feature aggregation. Caption features are reinjected during decoding via cross-attention to ensure that semantic guidance is preserved even when spatial similarity dominates neighbor selection. A lightweight decoder with refinement reduces patch artifacts and yields coherent high-resolution maps. Experiments on three benchmarks (LoveDA, Potsdam, and Vaihingen) show that the proposed framework achieves performance competitive with state-of-the-art methods while introducing a new language-guided paradigm for graph reasoning in RS image segmentation.