Large vision-language models (VLMs) can judge visual similarity, but their judgments are not directly available as compact image embeddings for efficient comparison. We study how to transfer these preferences into CLIP while retaining its image--text capabilities. We introduce ASK, a kernel-based steering method that l...
Sajjad Ghiasvand, Haniyeh Ehsani Oskouie, Sina Mansouri et al.· 0 citations
The concavity of the matrix-based conditional entropy functional is proved, which makes the resulting entropy-constrained projection a convex optimization problem, and a scalable mirror-descent algorithm for its implementation is developed.
These results suggest that in diversity-aware multi-armed bandits, e.g., for generative model selection, exploration can arise intrinsically from the objective's geometry, particularly for widely used metrics such as FID and Vendi where tight confidence bounds are difficult to construct.
B. Nia, F. Farnia· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.