GeoVP: A Unified Visual Prompting Framework for Multisource Remote-Sensing Image Understanding
Recent advances in prompt learning and multimodal large language models (MLLMs) have improved interactive image understanding. However, fine-grained remote-sensing (RS) interpretation remains challenging because text-only instructions are often insufficient to precisely specify regions of interest in complex scenes, and visual prompting methods developed for natural images generalize poorly to heterogeneous RS data. To address these challenges, GeoVP is proposed as a visual prompting MLLM for multisource RS image understanding. GeoVP supports point, box, and free-form prompts, enabling unified image-level and region-level understanding under different prompt granularities. It employs a hybrid vision encoder to extract multiscale semantic and structural features and a region-aware encoder to convert heterogeneous prompts into unified region representations. These cues are integrated with language instructions for fine-grained RS reasoning. A one-stage training strategy is adopted to improve cross-domain adaptation across natural-image and RS domains. In addition, an auxiliary Pixel-Level Localization Module provides qualitative mask-based visualization cues for prompted regions. GeoVP-650 K is constructed as a 654K-scale image–prompt–text triplet dataset covering optical, synthetic aperture radar, and infrared imagery. GeoVP achieves an average zero-shot classification accuracy of 79.05% on AID and an average cross-task classification accuracy of 87.55% on UCMerced. It also obtains an average semantic intersection over union of 98.25% on DIOR-RSVG under box prompts, demonstrating the effectiveness of explicit visual prompting for prompt-conditioned region understanding in multisource RS imagery.