TerraLogic, a benchmark for geospatial reasoning, is introduced and HieraPlan, a tool-augmented agent that organizes toolkits into functional hierarchies and performs fault-tolerant reasoning is proposed, providing a strong baseline with improved reasoning, cross-modal generalization, and error handling.
Abstract
Beyond perception, reasoning is essential in remote sensing for advanced interpretation, inference, and decision-making. Recent advances in large language models (LLMs) have enabled tool-augmented agents that leverage external tools to perform complex analytical tasks. However, existing studies in remote sensing primarily focus on perception-oriented tasks, leaving cognitive geospatial reasoning largely underexplored. To address this gap, we introduce TerraLogic, a benchmark for geospatial reasoning. TerraLogic comprises 545 scenario-driven, hierarchy-aware tasks, such as hazard vulnerability assessment, urban heat island analysis, and forest fragmentation dynamics, spanning optical, Synthetic Aperture Radar (SAR), and infrared (IR) imagery. It advances evaluation beyond recognition and monitoring toward cognitive-level geospatial analysis. To facilitate evaluation on TerraLogic, we further propose HieraPlan, a tool-augmented agent that organizes toolkits into functional hierarchies and performs fault-tolerant reasoning. HieraPlan enables structured abstraction, robust recovery from tool failures, and stable long-horizon planning. Extensive experiments demonstrate that current approaches struggle with hierarchical geospatial reasoning, while HieraPlan provides a strong baseline with improved reasoning, cross-modal generalization, and error handling. The dataset and agent code are publicly available at https://github.com/Ireliya/TerraLogic.
Geospatial vision-language models (VLMs) are increasingly expected to do more than assign scene labels or generate generic captions. In realistic Earth-observation and urban intelligence settings, useful multimodal systems must support fine-grained spatial reasoning, cross-view interpretation, grounded localization, and explicit understanding of change across time. Yet the current literature remains fragmented across remote sensing visual question answering, visual grounding, urban multi-view reasoning, and bi-temporal change captioning. As a result, claims about progress are often task-local, benchmark-specific, and difficult to compare. This paper reconstructs the field around a more defensible technical center: geospatial multimodal intelligence as the joint problem of spatial reasoning and temporal change understanding. Rather than presenting unverifiable new benchmark runs, we develop a rigorous review paper anchored in public benchmark evidence and a reference architecture for reproducible future work. We first formalize a task-centered problem definition that unifies image-level, region-level, cross-view, and bi-temporal reasoning. We then propose a reference GST-VLM architecture consisting of spatial encoding, temporal difference modeling, multimodal fusion, task-specific decoding, and reliability estimation. Next, we synthesize publicly reported evidence from representative datasets and benchmarks including RSVQA, EarthVQA, VRSBench, GeoChat, LEVIR-CD, LEVIR-CC, SECOND-CC, CHOICE, GEOBench-VLM, CityBench, and UrBench. The synthesis shows that recent models are improving rapidly but remain far from robust geospatial reasoning systems: on GEOBench-VLM, the best public model reported only 41.7% multiple-choice accuracy; on UrBench, even GPT-4o still trails human performance by an average 17.4 percentage points; and while specialized systems such as GeoReasoner, GeoChat, GeoLLaVA, and MModalCC outperform generic baselines on targeted tasks, their gains remain strongly benchmark-dependent. Based on this evidence, we identify the principal bottlenecks as benchmark fragmentation, weak temporal grounding, inadequate calibration, scarce cross-region validation, limited deployment reporting, and insufficient integration of geometry with language-conditioned reasoning. The paper concludes with a concrete research agenda for trustworthy geospatial VLMs that is centered on multi-temporal supervision, interactive change analysis, uncertainty-aware outputs, and evaluation protocols that measure not only accuracy but also transfer, calibration, and operational feasibility.
Abeer Hasshen Abdullah· International Journal of Inf...· 0 citations
Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challenging because EO workflows are constrained by sensing semantics, product dependencies, spatial and temporal compatibility, and parameter requirements. Existing agents often search a broad operation space for each query, while recent self-evolving systems do not fully organize heterogeneous EO trajectories into reusable knowledge across different decision levels. To solve this problem, we present GeoForge, a training-free, self-evolving framework that transforms completed trajectories into a structured nonparametric execution state. GeoForge constrains the operation space according to the sensing context, then retrieves a task-conditioned prior from three complementary memories. Workflow Graph Memory captures global operation order, Action-Level Experiences provide local corrections, and the Adapted Skill Standard Operating Procedure preserves procedural and data constraints. The retrieved prior guides tool execution, while current observations remain the basis of the final answer. After each task, a safety-gated distillation process converts grounded trajectories into reusable execution knowledge for future retrieval. This execution, distillation, and reuse loop improves planning without updating the backbone LLM. Experiments on multiple geospatial benchmarks demonstrate that GeoForge consistently improves both task accuracy and tool-use trajectory quality across diverse LLM backbones, while substantially reducing tool-planning and reasoning errors for most LLMs.
Xin Xiao, Jiang Zhong, Junnan Zhu et al.· 0 citations
Earth Observation (EO) analysis increasingly relies on large and heterogeneous satellite datasets, yet developing EO workflows often requires specialized expertise in data selection, geospatial programming, and cloud-based processing. Recent advances in Large Language Models (LLMs) offer new opportunities for natural-language interaction with EO systems, although challenges related to transparency, reproducibility, and domain-specific reasoning remain. This study presents SeaScope, an explainable AI framework that integrates LLMs, Retrieval-Augmented Generation (RAG), scientific knowledge retrieval, and Google Earth Engine (GEE) to transform natural-language requests into transparent and executable EO workflows. The framework combines knowledge retrieval, code generation, cloud execution, provenance tracking, and interactive visualization within a unified environment. A pilot implementation is demonstrated through maritime and coastal monitoring applications, including oil spill detection, vessel monitoring, water quality assessment, floating debris detection, and air quality analysis. Multiple state-of-the-art LLMs are evaluated under both RAG and non-RAG configurations using representative EO case studies. The results indicate substantial differences among model families and show that retrieval augmentation can significantly improve workflow generation quality and reliability for capable models, while providing more limited benefits for smaller models. The proposed framework demonstrates the potential of explainable AI agents to support transparent, reproducible, and scalable EO analysis.
Geotechnical hazards such as landslides, subsidence, slope instability, and infrastructure deformation threaten rapidly urbanising and environmentally stressed regions worldwide, intensifying the need for scalable and intelligent monitoring systems capable of continuously observing complex Earth surface dynamics. Although multisensor remote sensing fusion has substantially expanded the observational capabilities of modern geotechnical monitoring through the integration of Synthetic Aperture Radar (SAR), optical imagery, Light Detection and Ranging (LiDAR), and environmental data, existing fusion pipelines remain subject to several well-documented constraints, including weak semantic alignment, limited temporal reasoning, and poor transferability across heterogeneous environmental conditions. This review synthesises the emerging transition from conventional sensor-centric fusion toward intelligent geospatial monitoring architectures centred on deep multimodal representation learning, transformer-based temporal reasoning, self-supervised learning, and geospatial foundation models. Particular emphasis is placed on how recent architectures are designed to better preserve coherent spatial, temporal, and contextual environmental relationships within unified latent representation spaces rather than through downstream handcrafted integration. The review further examines the growing role of multimodal transformers, masked autoencoders, contrastive learning, and large-scale geospatial foundation models in enabling scalable environmental reasoning, adaptive multimodal learning, and transferable geospatial intelligence across sensing modalities and geographic domains. Finally, remaining challenges involving uncertainty, explainability, computational scalability, and environmental generalisation are discussed alongside future research directions involving continual learning, physics-aware artificial intelligence, and autonomous geotechnical monitoring systems. Together, the reviewed literature suggests that multimodal Earth observation is evolving from passive environmental sensing toward adaptive geospatial intelligence systems capable of scalable hazard reasoning and autonomous environmental understanding.
Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason across time and space. However, existing studies lack a systematic evaluation that dissects these distinct competencies. To fill this gap, we introduce ChronoBench, a multidimensional benchmark that decomposes this task into four progressive cognitive levels (i.e., Land Cover Perception, Temporal Recognition, Long-Term Memory, and Spatio-Temporal Reasoning). The ChronoBench comprises 12 sub-tasks and 17,689 rigorously validated QA (Question-Answer) pairs. Extensive evaluations reveal that mainstream MLLMs fall drastically behind human experts, with Long-Term Memory emerging as the most critical bottleneck. Motivated by this finding, we further propose GeoChrono, an MLLM with enhanced capabilities for tracing, memorizing, and reasoning about long-term geographic evolution. Leveraging the physical prior that geographic parcels remain spatially fixed while their semantics evolve, we design a Temporal Trajectory Encoder~(TempEnc) that constructs per-location temporal trajectories for dedicated land cover evolution modeling, and we introduce a Coarse-to-Fine Token Compressor~(C2FComp) that adaptively preserves dynamic regions while compressing the static background. To support training, we also construct ChronoInstruct, a 104K-sample instruction-tuning dataset spanning all competency levels for training. GeoChrono achieves state-of-the-art performance on ChronoBench, surpassing the leading commercial MLLMs by over 20%, while C2FComp reduces visual tokens by over 56% while retaining GeoChrono's 94.6% performance. The code and data will be available at https://github.com/IntelliSensing/GeoChrono
Yujie Li, Jiancheng Pan, Zhiwei Wei et al.· 0 citations
Geospatial technology has evolved from a set of specialised mapping tools into an integrated decision infrastructure that links Earth observation, geographic information systems, satellite positioning, uncrewed aerial systems, volunteered geographic information, cloud computing and geospatial artificial intelligence. This critical narrative review examines how that integration is changing evidence production and decision support across agriculture, environmental and biodiversity monitoring, disaster risk management, urban and transport planning, and public health. Literature published primarily from 2000 to 9 June 2026 was selected through live web-based scholarly discovery, DOI and bibliographic verification, citation chaining and targeted searches of accessible scholarly records. The evidence indicates that geospatial systems are most valuable when they combine complementary observations across scales rather than relying on a single sensor, platform or algorithm. Their strongest contributions are spatially explicit monitoring, prioritisation, scenario analysis and repeated observation, while the weakest parts of many workflows remain ground-reference quality, uncertainty propagation, transferability, interoperability and evaluation against operational outcomes. Cloud platforms and deep learning have expanded computational reach, but they can also conceal provenance, amplify geographic bias and encourage benchmark-driven optimisation that does not translate reliably across places. Volunteered data and urban digital twins extend participation and real-time representation, yet introduce uneven coverage, privacy, governance and accountability concerns. The review argues that future progress depends less on incremental accuracy gains than on reproducible multi-source workflows, explicit uncertainty, trustworthy and explainable geospatial artificial intelligence, privacy-preserving governance, interoperable standards and evaluation in the institutions that ultimately use the evidence. Geospatial technology should therefore be judged not only by spatial resolution or predictive performance, but by whether it produces defensible, equitable and actionable knowledge across heterogeneous real-world settings.
Yashvardhan Singh, Divya Singh, Sakshi Shukla et al.· Advances in Research· 0 citations