Experiments with a state-of-the-art CVL model on OpenCVL show that incorporating noisy in-the-wild data consistently improves performance on clean test sets, suggesting a promising direction for scaling CVL with diverse real-world imagery.
Abstract
Fine-grained Cross-View Localization (CVL) estimates the precise position and orientation of a ground-level image by aligning it with geo-referenced aerial imagery, offering a scalable alternative to Global Navigation Satellite Systems (GNSS) in challenging urban environments. Existing datasets rely on data collected with high-end sensor suites, which inherently limit image diversity and scalability. While in-the-wild images are abundant, their noisy geo-tags make them unsuitable for reliable evaluation. To bridge this gap, we introduce OpenCVL, a large-scale, diverse, and open dataset containing 617,388 ground-aerial image pairs spanning 41 cities across four European countries. All images are sourced from permissive platforms, ensuring long-term accessibility and supporting open and reproducible research. The training set combines images captured with high-end sensors with diverse in-the-wild imagery. We further develop a data curation framework that filters and corrects pose annotations to construct reliable in-the-wild evaluation data. In addition, OpenCVL includes dedicated cross-area and snowy test sets to assess generalization and robustness. Experiments with a state-of-the-art CVL model on OpenCVL show that incorporating noisy in-the-wild data consistently improves performance on clean test sets, suggesting a promising direction for scaling CVL with diverse real-world imagery.
Remote sensing object detection remains a highly challenging task due to drastic variations in target scale, dense distributions of small objects, and complex background interference. Although existing detectors have achieved notable progress, they heavily rely on large-scale supervised pretraining and often suffer from substantial domain gaps and expensive annotation costs. In recent years, vision foundation models (VFMs) have demonstrated strong general-purpose representation capability, yet their potential in remote sensing imagery has not been fully explored. To bridge this gap, we propose SKYDET, an end-to-end robust object detection framework that explicitly migrates the billion-parameter DINOv3 model to the aerial domain. To effectively bridge the domain gap and prevent representation manifold degradation, we freeze the pretrained foundation model and propose a semantic guiding adapter (SGA) that acts as a precise semantic filter to suppress irrelevant background clutter. In addition, to address feature misalignment and semantic ambiguity during cross-scale fusion, we introduce a cross-fused encoder (CFE), whose core component is the reciprocal guidance module (RGM). The RGM establishes a reciprocal enhancement mechanism that enables spatial structure and channel semantics to guide each other, thereby effectively suppressing background noise and strengthening responses to small objects. Extensive experiments on three challenging benchmark datasets—DOTA-v1.0, AI-TOD, and NWPU VHR-10—demonstrate highly competitive performance. Specifically, SKYDET-C achieves state-of-the-art (SOTA) detection precision (with $AP_{50}$ scores of 72.6%, 56.1%, and 95.6%, respectively), while SKYDET-T exhibits superior advantages in high-precision localization tasks. This validates the effectiveness of transferring VFM knowledge to remote sensing tasks. Our code is available at https://github.com/zhangyyy-ai-rs/SKYDET
Yao-Dong Zhang, Wei Guo, Boxiang Xie et al.· IEEE Transactions on Geoscie...· 0 citations
Cross-view geo-localization (CVGL) retrieves geo-tagged satellite imagery for a ground-view query. Most systems exhaustively search a flat, fixed-resolution gallery, incurring high cost over large areas and adapting poorly to satellite resolution changes. Autoregressive coarse-to-fine alternatives reduce comparisons but bind later predictions to earlier decisions and a predefined hierarchy. We introduce GeoMoE, a sparse mixture-of-experts dual encoder that decouples global multi-scale representation learning from local hierarchical search. Global multi-scale supervision and content-adaptive routing map ground and satellite images across resolutions into a globally comparable embedding space. At inference, each image is encoded once, and probabilistic beam search follows parent--child links to score a small candidate subset. Later levels reuse these descriptors rather than features generated by preceding levels, limiting feature-level error propagation and hierarchy coupling. We further introduce VIGOR-M, a four-city benchmark with an explicit parent--child satellite hierarchy and held-out half-step galleries for single-resolution, cross-resolution, and hierarchical evaluation. GeoMoE achieves 95.78% R@40m on Just Zoom In, 2.77 percentage points above the previous best, and 62.39% R@1 on VIGOR-M. The latter requires 0.885 MMAC/query for descriptor matching, 5.27% of an exhaustive L3 scan, while exceeding the strongest exhaustive baseline by 3.12 percentage points in R@1. One model trained on L1, L2, and L3 also outperforms a matched dense control across all six galleries and transfers to three withheld resolutions. By decoupling globally trained embeddings from local hierarchical search, GeoMoE jointly improves localization accuracy, search efficiency, and cross-resolution transfer.
Ruijie Fan, Junyan Ye, Qiyuan Zhu et al.· 0 citations
This work introduces OpenAqua, the first large-scale fine-grained dataset dedicated to open underwater visual tasks, and establishes a comprehensive benchmark suite that encompasses not only standard object detection and instance segmentation tasks but also pioneers an underwater open-vocabulary object detection benchmark.
Linxuan Luo, Pan Mu, Cong Bai· Proceedings of the 32nd ACM...· 0 citations
Satellite object detection is challenged by small targets and wide-format scenes that lose detail under standard square-input resizing. We introduce SkySeaLand, a public dataset of 1,307 high-resolution satellite images and 19,101 verified bounding boxes across airplane, boat, car, and ship classes in terrestrial and maritime scenes. Native COCO and YOLO annotations are provided. The collection is dominated by large source images and wide scene geometry: 84.5 percent exceed 3,836 pixels on the longest side and 73.1 percent are near a 3:1 aspect ratio. We evaluate twelve detectors from the YOLO, RT-DETR, DETR, and Faster R-CNN families using a common split and COCO metrics. The tested YOLO and RT-DETR variants obtain 84.4--88.2 mAP50, with no consistent accuracy gain from larger parameter counts under the reported model-specific recipes. We also report SkyDet, a 1.22 M parameter anchor-free baseline that obtains 60.5 mAP50 and 24.32 mAP50-95 in a 4.90 MB footprint, with 13.74 ms latency (72.8 FPS) on a Tesla T4. SkySeaLand provides a compact benchmark for mixed land--maritime transportation detection, while SkyDet establishes a documented low-footprint reference rather than a state-of-the-art accuracy claim.
Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recently, the recent release of DINOv3 brings DINO.txt, which equips the standalone DINO backbone with image-text contrastive learning and thus opens up the possibility of open-vocabulary segmentation. We propose DinoSplat-OV, a training-free framework that adapts DINOv3 to remote sensing without fine-tuning or additional pretraining. Targeting the dense distribution, multi-scale nature, and large size of remote sensing imagery, we design two core modules. Its Text-aware Laplacian Propagation module de-noises patch-level predictions by combining textual semantic affinities with local visual similarity, improving regional consistency while preserving boundaries. Its Gaussian Splatting Upsampling module reconstructs pixel-level features through RGB-guided anisotropic aggregation and test-time optimization. A global-anchor sliding-window strategy further supports large-scale imagery. Experiments on UDD5, DOTA, and LoveDA demonstrate competitive or superior performance over existing training-free methods, effectively filling the gap of DINO-series models in training-free open-vocabulary segmentation and providing a viable new path for further advances in this direction.
Changhao Zhao, Haoxiang Li, Yuke Li et al.· 0 citations