Skip to content

STAR: Spatio-Temporal Alignment and Refinement for Cross-View Geo-Localization

2026 · IEEE Transactions on Geoscience and Remote Sensing · Vol 64, pp. 4706213-4706213 · 0 citations · 44 references

Abstract

Cross-view geo-localization (CVGL) aims to match images captured from different viewpoints, such as drone and satellite imagery. Existing methods primarily focus on single-image matching, overlooking the potential of leveraging temporal information from drone image sequences. To address this, we propose Spatio-Temporal Alignment and Refinement (STAR), a novel framework that effectively leverages the spatio-temporal dependencies of drone sequences for effective matching with satellite images. Notably, this is among the first attempts to explore drone-sequence-based matching in CVGL, further improving retrieval performance through spatio-temporal consistency. Specifically, we design the Video Vision Transformer (ViViT) Sequence Alignment Module (VSAM) to effectively extract and align sequential features and introduce a Pseudo-Temporal Compensation (PTC) strategy to prevent the temporal modeling branch from degenerating when processing static satellite inputs, maintaining architectural symmetry across dynamic and static views. In addition, we design the Dynamic Background Partitioning Module (DBPM) to adaptively segment features and improve foreground–background distinction. Extensive experiments on the University-1652 and SUES-200 datasets demonstrate that our method significantly outperforms state-of-the-art approaches, highlighting the benefits of incorporating sequence information in CVGL.

View source