Skip to content

TecoPrompt: Temporal-Conservative Prompt Learning for Vision-Language Models

Sep 2026 · 0 citations · 52 references
Computer Science

TL;DR

TecoPrompt, a closed-loop robust prompt-learning framework that revisits optimal transport pseudo-labeling from a temporal perspective, employs an entropic OT plan in the CLIP semantic space to obtain globally consistent label candidates.

Abstract

Prompt learning adapts vision-language models, such as CLIP, by adjusting a small set of context tokens. However, under few-shot supervision, even moderate label noise can disrupt prompt optimization. To address this issue, we propose TecoPrompt, a closed-loop robust prompt-learning framework that revisits optimal transport (OT) pseudo-labeling from a temporal perspective. TecoPrompt employs an entropic OT plan in the CLIP semantic space to obtain globally consistent label candidates. It verifies the reliability of these candidates by examining trajectory stability: a noisy label is only rewritten if the OT candidate remains unchanged within a K-epoch temporal stability window and passes a confidence gate based on Exponential Moving Average (EMA). This approach helps reduce confirmation bias. The rewritten labels are then integrated back into prompt training using a tri-group objective that includes three loss functions aligned with clean, mid, and noisy subsets. Experiments on seven datasets with synthetic symmetric and asymmetric noise, as well as Food101N, demonstrate significant performance improvements. For example, on the OxfordPets dataset, with 50% asymmetric noise, TecoPrompt achieves an accuracy of 0.843, up from 0.775.

View source

Similar papers

#small language model Preprint Sep 2026

ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

ProCAP is proposed, a probabilistic cross-attentive prompt learning framework that improves cross-modal interaction and training stability without updating any CLIP weights: it learns both visual and textual prompt tokens and links them through stacked bidirectional multi-head cross-attention so the two branches refine...

Hiwa Azeez Abbas, Fatemeh Daneshfar, M. Abdar · 0 citations
Conference Open access Sep 2026

CoDA: Co-adaptive Dual-path Alignment for Vision-Language Models

In CoDA, a new adaptation framework that explicitly disentangles and coordinates cross-modal semantic alignment and intra-modal structural consistency is proposed, and it is shown that CoDA outperforms state-of-the-art parameter-efficient methods, particularly under few-shot learning and distribution-shift scenarios.

Yi Zhang, Rui Zhu, Chan-Ni Li et al. · 0 citations
Preprint Aug 2026

PPOM: Marginalizing Patch-Grid Phase for CLIP-Based Generalizable Vision-Language Prompt Tuning

Prompt tuning adapts CLIP-based vision-language models with few trainable parameters, yet its predictions remain sensitive to the spatial sampling imposed by a frozen vision transformer. In particular, non-overlapping patch tokenization makes predictions depend on the alignment (phase) between image and the patch latti...

Liang Wang, Haoyang Li, Chao Wang et al. · 0 citations
#machine learning Preprint Sep 2026

Text Capability Loss in Vision-Language Adaptation: An Attention-Sink Diagnosis

It is found that post-pretraining QK-RMSNorm injection fails to reproduce the protection of native QK-RMSNorm, while several off-the-shelf weight-merging settings fail to recover the lost capability after VL training, highlighting the value of screening backbones with Sink Strength before VL training and narrow the int...

Minsik Choi, Geewook Kim, Young Geun Kim · 0 citations
#artificial intelligence Preprint Sep 2026

PETR: Prompt Ensembling with Training-free Routing for Vision-Language Models

Prompt learning efficiently adapts vision-language models (VLMs) to downstream tasks, but gains on seen classes often come at the expense of generalization to unseen classes. To address this limitation, we propose prompt ensembling with training-free routing (PETR), whose key innovation is a carefully designed dual-pro...

Wei-Han Cai, Hao Tan, Xin-Ping Gao et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.