This work forms Progressive Cross-view Video Geo-localization (PCVG) as a deployment-oriented extension and evaluation protocol of CVG, enabling localization under varying temporal budgets, prefix-based inference, random-start evaluation, and long-range localization with interruptions.
Abstract
Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. However, CVG approaches rely on fixed-length inputs and post-hoc refinement, hindering online-oriented localization under partial or dynamic observations. In this work, we formulate Progressive Cross-view Video Geo-localization (PCVG) as a deployment-oriented extension and evaluation protocol of CVG, enabling localization under varying temporal budgets, prefix-based inference, random-start evaluation, and long-range localization with interruptions. To explore PCVG, we introduce X$^2$Localizer, a cross-grained alignment framework that jointly supervises global prefix-to-aerial retrieval and token-aggregated frame--aerial-tile matching with a budget-dependent asymmetric objective. Furthermore, we introduce a Sliding-Window Re-Localization (SWRL) strategy that dynamically refreshes candidate regions for failure recovery and long-range deployment without full-sequence reprocessing. Extensive experiments show that X$^2$Localizer preserves conventional full-video performance, with marginal gains of +0.1 Recall@1 and +0.3 Recall@10, while substantially improving early localization. In the challenging single-frame setting, X$^2$Localizer improves coarse retrieval by +4.7 Recall@1 and +11.5 Recall@10 over the previous state-of-the-art method. With SWRL, our approach further enables robust progressive localization under random-start and long-distance scenarios, narrowing the gap between benchmark evaluation and real-world deployment.
Cross-view geo-localization (CVGL) aims to match images captured from different viewpoints, such as drone and satellite imagery. Existing methods primarily focus on single-image matching, overlooking the potential of leveraging temporal information from drone image sequences. To address this, we propose Spatio-Temporal...
Ziqian Mo, Yu-Xi Sun, Sen Jia et al.· IEEE Transactions on Geoscie...· 0 citations
Cross-view geo-localization (CVGL) is a critical task that determines the geographic position of a query image via retrieving its corresponding counterparts across heterogeneous visual domains, such as ground-level, drone, and satellite views. Despite significant progress, recent existing approaches primarily focus on...
Guan-Bo Wang, Xu-Lei Shi, Xin Wang et al.· IEEE Geoscience and Remote S...· 0 citations
Cross-view object geo-localization (CVOGL) typically assumes a fixed query image, overlooking the ability of mobile agents to actively acquire more informative observations. To address this limitation, we introduce Active Cross-View Object Geo-Localization (ActiveGeo), where an agent sequentially selects new viewpoints...
Shun-Yu Yao, Xiao-Han Zhang, Zhuo-Ran Yang et al.· 0 citations
Cross-View Geo-Localization (CVGL) with OpenStreetMap (OSM) performs well in structure-rich urban environments but collapses in feature-sparse scenes such as rural roads. To study this failure mode, in this work, we introduce CV-FSS, a benchmark that pairs sequential panoramas from five rural regions with aligned OSM m...
Junwei Zheng, Yunyi Huang, Ruize Dai et al.· 1 citation
Cross-view geo-localization (CVGL) is a critical technology used in unmanned aerial vehicles (UAVs) and widely applied in navigation and target localization tasks. However, owing to the extreme perspective disparity between UAV oblique views and satellite vertical views, CVGL still involves significant challenges, incl...
Xiaojia Yan, Zhang-Song Shi, Shiyan Sun et al.· Drones· 0 citations
X-GeoP2P, a coarse-to-fine localization pipeline that combines adapted versions of GAReT and X-VGGT, is presented, a coarse-to-fine localization pipeline that combines adapted versions of GAReT and X-VGGT.
Wyatt Petula, Soumik Ghosh, Qingyang Liu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.