Skip to content
Open access

Semisupervised Adaptation of Vision-Language Models for Image Classification

Aug 2026 · IEEE Geoscience and Remote Sensing Letters · Vol 23, pp. 4014105-4014105 · 0 citations · 16 references
Computer Science

TL;DR

Results on the UC Merced (UCM) and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches and provides a viable solution for adapting VLMs to the remote sensing domain with minimal human intervention.

Abstract

Vision-language models (VLMs), like contrastive language-image pretraining (CLIP), have shown significant potential in handling natural images, yet their performance is often limited by the distinct characteristics of satellite imagery. While parameter-efficient adaptation techniques exist, their efficacy is frequently limited by the scarcity of annotated samples. In this letter, we propose self-evolutionary CLIP (SE-CLIP), a semisupervised framework designed for recursive label mining in scene classification. The approach follows a dual-phase pipeline, where an initial warmup on a few annotated seeds is followed by a recursive discovery phase that iteratively identifies high-confidence samples from unlabeled pools. To maintain the integrity of the evolving support set, we employ a class-balanced selection strategy that prevents the model from being dominated by easily learned categories. Results on the UC Merced (UCM) and NWPU benchmarks indicate that SE-CLIP significantly outperforms existing semi-supervised approaches. The framework provides a viable solution for adapting VLMs to the remote sensing (RS) domain with minimal human intervention.

Read PDF

Similar papers

2026

Toward Zero-Forgetting: A Training-Free Multimodal Framework for Remote Sensing Class-Incremental Learning

Existing class-incremental learning (CIL) methods for remote sensing (RS) scene classification often tend to be training-intensive or rely on static visual features that may inadequately capture the complex interclass similarity and intraclass diversity inherent in RS imagery. Moreover, directly reusing features from m...

Wen-Liang Du, Ji-Cun He, Jia-Qi Zhao et al. · 0 citations
Preprint Aug 2026

MLLMCLIP: Feature-Level Distillation of MLLM for Robust Vision-Language Representations

Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capabil...

Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao et al. · 0 citations
Preprint Aug 2026

UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models

UC-VLM is a unified multi-stage binary-supervised framework that consistently reuses the same authenticity labels for visual adaptation and label-conditioned text generation, while leveraging automatically optimized instructions to reduce prompt sensitivity without requiring human-written rationales or hand-crafted pro...

Lei Tan, Shuwei Li, Mohan S. Kankanhalli et al. · 0 citations
Preprint Aug 2026

Clustering and Token Denoising for Faster and More Robust VLMs

This study introduces ClustRS, a two-part, training-free algorithm for robust token pruning, demonstrating a simple yet powerful alternative to both score-only and diversity-only pruning rules, paving the way for compute-efficient and noise-resilient VLM deployment.

Baptiste Rossigneux, Inna Kucher, Vincent Lorrain et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.