Aug 2026· Applied Sciences· 0 citations· 19 references
TL;DR
GeoGATE is introduced, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware temporal reasoning that associates adaptive slicing most strongly with localization, retrieval with language and VQA, and language model adaptation with all reported tasks.
Abstract
High-resolution remote sensing understanding requires models to preserve small spatial evidence, account for acquisition-dependent appearance, and separate genuine geographic change from nuisance variation. We introduce GeoGATE, a geo-sensor-guided framework that combines typed acquisition conditioning, budget-constrained adaptive token acquisition, metadata-compatible evidence retrieval, and reliability-aware temporal reasoning. LoRA adaptation and NF4 quantization support efficient training and deployment. On the VRSBench test split, GeoGATE reaches 53.4 BLEU-1, 36.8 BLEU-2, 18.2 BLEU-4, 56.4 Acc@0.5, 82.3 VQA, 25.1 METEOR, and 42.6 ROUGE-L, outperforming the controlled GeoGATE (Base) configuration across captioning, question answering, and grounding. Component ablations associate adaptive slicing most strongly with localization, retrieval with language and VQA, and language model adaptation with all reported tasks. NF4 reduces measured video memory from 24.5 GiB to 7.2 GiB with only minor metric changes. These experiments support the single-image language and grounding components. Dedicated cross-sensor and bi-temporal benchmarks are not reported; the corresponding modules are therefore presented as architectural extensions rather than validated performance claims.
It is shown that higher resolution is not universally beneficial and, once sufficient granularity is reached, token quality matters more than token quantity, and a plug-and-play framework that unifies task-adaptive token granularity allocation with holistic geospatial token importance modulation is presented.
Kexin Ma, Jing Xiao, Bowen Xing et al.· 0 citations
Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information from such data remains challenging. In particular, fine-grained regional geolocalization is challenging because nearby areas often share coarse geographic cues. We study...
Chang-Le Lee, Yeonsoo Park, Abdullah Alfarrarjeh et al.· 0 citations
AlignJEPA provides a parameter-efficient route for aligning Earth observation foundation models with language by using a pretrained AnySat visual encoder and a RemoteCLIP text encoder while training only a lightweight predictive alignment network.
Md Aminur Hossain, Omkumar Vaghasiya, R. Dwivedi et al.· 0 citations
Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires models to resolve small visual evidence within extremely large images. Existing approaches typically rely on token pruning, visual search, or tool-augmented reasoning at inference time. We instead investigate whether the benefit of zoom...
Chengjie Jiang, Yun-Qi Zhou, Jia-Feng Yan et al.· 0 citations
Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short seque...
Yupan Ding, Jing Xiao, Zhenyuan Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.