UniGeo, a unified multimodal large language model (MLLM) for text-guided drone geo-localization, is proposed and a multi-stage training strategy that progressively learns geo-semantic understanding, cross-view generation, and candidate verification, improving adaptation to text-guided geo-localization is introduced.
Abstract
Text-guided drone geo-localization aims to identify a target region in a large-scale image gallery from a natural-language description. Existing methods mainly formulate this task as direct matching between an open-ended text query and candidate images. However, incomplete queries and highly similar candidates often make global cross-modal matching insufficient for reliable fine-grained localization. We propose UniGeo, a unified multimodal large language model (MLLM) for text-guided drone geo-localization. Built on a shared vision-language framework, UniGeo jointly supports geo-semantic understanding, cross-view semantic generation, and candidate-level verification. Specifically, it establishes stable correspondences among local scene elements, spatial relations, and language descriptions through geo-semantic learning, and further models semantic mappings between drone and satellite views through cross-view generation. Based on these capabilities, a plug-and-play verification module performs fine-grained discrimination among highly confusable candidates. We further introduce a multi-stage training strategy that progressively learns geo-semantic understanding, cross-view generation, and candidate verification, improving adaptation to text-guided geo-localization. Experiments demonstrate consistent improvements across multiple retrieval backbones. On GeoText-1652, UniGeo improves R@10 and mAP by 13.59 and 2.83 percentage points, respectively, validating its effectiveness for fine-grained text-guided drone geo-localization.
The proposed Cross-Modal Adaptive Token Selection and Alignment Network (CATSANet), a CLIP-based framework tailored for TI-ReID, achieves competitive performance in terms of Rank-k accuracy and mAP, demonstrating the effectiveness of fine-grained alignment and ranking refinement across datasets.
Dongbin Chen, Junjie Li, Hao Xu et al.· Pattern Analysis and Applica...· 0 citations
Recent advances in remote sensing (RS) vision-language foundation models (VLFMs) rely heavily on large-scale paired image–text data. However, existing dataset construction methods mainly emphasize data scale while overlooking a more fundamental limitation: the lack of structured and comprehensive semantic representatio...
Yi-Guo He, Jun-Jie Zhu, Jun Wang et al.· IEEE Transactions on Geoscie...· 0 citations
Vision-Language Pre-training (VLP) aims to convert vision-language tasks into image-text matching challenges, necessitating robust interaction between visual and textual modalities. However, current research often faces a trade-off: methods pursuing fine-grained alignment typically rely on computationally expensive obj...
Jun-Lin Jiang, D. Vuković, Jie Cao et al.· ACM Transactions on Multimed...· 0 citations
Scene Text Retrieval (STR) aims to search images containing a given textual query within large-scale image collections. However, existing approaches are fundamentally constrained in two ways: 1) they are evaluated on narrow benchmarks that focus primarily on natural scenes; and 2) they rely on either error-prone multi-...
Tong-Kun Guan, Yu-Tong Cai, Hao-Cheng Wang et al.· IEEE Transactions on Image P...· 0 citations
Point-of-interest (POI) localization matches user-provided storefront close-ups to the same shops in wide, geo-tagged vehicle-mounted street views. POIs may change while the surrounding scene stays similar, so scene-level recognition alone cannot establish POI identity. Differences in target scale and capture domains f...
Lu-Hua Han, Xi-Ting Sun, Hao Wang et al.· 0 citations
Cross-view object geo-localization (CVOGL) locates a target in satellite imagery using drone or street-view queries. Existing methods train separate detectors for each viewpoint, leading to parameter redundancy and impeding cross-view knowledge sharing. Moreover, top-ranked satellite candidates are often visually simil...
Xuezhi Fan, Ming Qi, Zhu Han et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.