Large vision-language models (VLMs) can judge visual similarity, but their judgments are not directly available as compact image embeddings for efficient comparison. We study how to transfer these preferences into CLIP while retaining its image--text capabilities. We introduce ASK, a kernel-based steering method that l...
Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Sina Mansouri et al.· 0 citations
Across seven few-shot benchmarks and two CLIP backbones, low-rank prompts match or improve dense CoOp at far fewer parameters, with the clearest gains on low-shot base-to-new generalization.
This work proposes ZOMP (Zeroth-Order Multimodal Prompt tuning), a query-efficient, fully forward-only method that tunes deep prompts in both the vision and text branches of a frozen CLIP model using simultaneous perturbation stochastic approximation.
Sajjad Ghiasvand, Yifan Yang, Mahnoosh Alizadeh et al.· 1 citation
It is argued that reference-based similarity rewards a fluent, comprehensive critique style rather than the selectivity and specificity of human critique, and that reference-based similarity gives a misleading picture.
REALM is proposed, which jointly learns the model parameters and a scalar expertise value for each annotator, entirely unsupervised and requiring nothing beyond annotator identity, and extends to multiple tasks via a learned expertise matrix.
Sajjad Ghiasvand, M. Beliaev, Mahnoosh Alizadeh et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.