Aug 2026· IEEE Geoscience and Remote Sensing Letters· Vol 23, pp. 4014105-4014105· 0 citations· 16 references
Computer Science
TL;DR
Results on the UC Merced (UCM) and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches and provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.
Abstract
Vision-language models (VLMs), like contrastive language-image pretraining (CLIP), have shown significant potential in handling natural images, yet their performance is often limited by the distinct characteristics of satellite imagery. While parameter-efficient adaptation techniques exist, their efficacy is frequently limited by the scarcity of annotated samples. In this letter, we propose self-evolutionary CLIP (SE-CLIP), a semisupervised framework designed for recursive label mining in scene classification. The approach follows a dual-phase pipeline, where an initial warmup on a few annotated seeds is followed by a recursive discovery phase that iteratively identifies high-confidence samples from unlabeled pools. To maintain the integrity of the evolving support set, we employ a class-balanced selection strategy that prevents the model from being dominated by easily learned categories. Results on the UC Merced (UCM) and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches. The framework provides a viable solution for adapting VLMs to the remote sensing (RS) domain with minimal human intervention.
Across cross-generator, post-processing, and in-the-wild benchmarks, PE-SPC surpasses the previous DINOv3 baseline and achieves new state-of-the-art results.
Wei-Han Cai, Hao Tan, Zichang Tan et al.· 0 citations
Existing class-incremental learning (CIL) methods for remote sensing (RS) scene classification often tend to be training-intensive or rely on static visual features that may inadequately capture the complex interclass similarity and intraclass diversity inherent in RS imagery. Moreover, directly reusing features from m...
Wen-Liang Du, Ji-Cun He, Jia-Qi Zhao et al.· IEEE Transactions on Geoscie...· 0 citations
Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capabil...
Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao et al.· 0 citations
UC-VLM is a unified multi-stage binary-supervised framework that consistently reuses the same authenticity labels for visual adaptation and label-conditioned text generation, while leveraging automatically optimized instructions to reduce prompt sensitivity without requiring human-written rationales or hand-crafted pro...
Lei Tan, Shuwei Li, Mohan S. Kankanhalli et al.· 0 citations
This study introduces ClustRS, a two-part, training-free algorithm for robust token pruning, demonstrating a simple yet powerful alternative to both score-only and diversity-only pruning rules, paving the way for compute-efficient and noise-resilient VLM deployment.
Baptiste Rossigneux, Inna Kucher, Vincent Lorrain et al.· 0 citations