Skip to content
Open access

GeoGATE: Geo-Sensor-Guided Adaptive Token and Evidence Reasoning for High-Resolution Remote Sensing Image Understanding

Aug 2026 · Applied Sciences · 0 citations · 19 references

TL;DR

GeoGATE is introduced, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware temporal reasoning that associates adaptive slicing most strongly with localization, retrieval with language and VQA, and language model adaptation with all reported tasks.

Abstract

High-resolution remote sensing understanding requires models to preserve small spatial evidence, account for acquisition-dependent appearance, and separate genuine geographic change from nuisance variation. We introduce GeoGATE, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware temporal reasoning. LoRA adaptation and NF4 quantization support efficient training and deployment. On the VRSBench test split, GeoGATE reaches 53.4 BLEU-1, 36.8 BLEU-2, 18.2 BLEU-4, 56.4 Acc@0.5, 82.3 VQA, 25.1 METEOR, and 42.6 ROUGE-L, outperforming the controlled GeoGATE (Base) configuration across captioning, question answering, and grounding. Component ablations associate adaptive slicing most strongly with localization, retrieval with language and VQA, and language model adaptation with all reported tasks. NF4 reduces measured video memory from 24.5 GiB to 7.2 GiB with only minor metric changes. These experiments support the single-image language and grounding components. Dedicated cross-sensor and bi-temporal benchmarks are not reported; the corresponding modules are therefore presented as architectural extensions rather than validated performance claims.

Read PDF

Similar papers

Preprint Aug 2026

SA-GEM: Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning for Efficient Remote Sensing Large Vision-Language Models

It is shown that higher resolution is not universally beneficial and, once sufficient granularity is reached, token quality matters more than token quantity, and a plug-and-play framework that unifies task-adaptive token granularity allocation with holistic geospatial token importance modulation is presented.

Kexin Ma, Jing Xiao, Bowen Xing et al. · 0 citations
Preprint Aug 2026

What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation

Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information from such data remains challenging. In particular, fine-grained regional geolocalization is challenging because nearby areas often share coarse geographic cues. We study...

Chang-Le Lee, Yeonsoo Park, Abdullah Alfarrarjeh et al. · 0 citations
Preprint Aug 2026

AlignJEPA: Predictive Vision-Language Alignment for Remote Sensing Foundation Models

AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language by using a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network.

Md Aminur Hossain, Omkumar Vaghasiya, R. Dwivedi et al. · 0 citations
Preprint Sep 2026

RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA

Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom...

Chengjie Jiang, Yun-Qi Zhou, Jia-Feng Yan et al. · 0 citations
Preprint Aug 2026

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short seque...

Yupan Ding, Jing Xiao, Zhenyuan Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.