Skip to content
Preprint

Low-Rank Prompt Learning for Vision-Language Models with Fixed-Token Bases

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

Across seven few-shot benchmarks and two CLIP backbones, low-rank prompts match or improve dense CoOp at far fewer parameters, with the clearest gains on low-shot base-to-new generalization.

Abstract

Prompt learning adapts CLIP to downstream recognition by replacing hand-written templates with learned continuous context vectors, which in Context Optimization (CoOp) form a dense prompt matrix $\mathbf{P}\in\mathbb{R}^{m\times d}$ trained from only a few examples per class. We study whether this matrix is over-parameterized by factorizing it as $\mathbf{P}=\mathbf{B}\mathbf{A}$, which cuts the trainable prompt parameters from $md$ to $r(m+d)$, and to $rd$ once the token-side factor $\mathbf{B}$ is fixed. Across seven few-shot benchmarks and two CLIP backbones, low-rank prompts match or improve dense CoOp at far fewer parameters, with the clearest gains on low-shot base-to-new generalization. We then find that the token-side factor need not be learned at all: fixing $\mathbf{B}$ to a Gaussian, orthogonal, SVD-derived, or even random basis and training only the embedding-side factor $\mathbf{A}$ stays on par with the fully trainable factorization, and a source-trained $\mathbf{B}$ offers no advantage over a random one. A prompt-factor asymmetry and a local update-space dimension gap show why fixing $\mathbf{B}$ is far less restrictive than fixing $\mathbf{A}$, and a smoothness-only guarantee certifies that optimizing $\mathbf{A}$ over a fixed $\mathbf{B}$ converges. In the CLIP prompt setting, the embedding-side coefficients carry the adaptation while the token basis can simply be fixed.

View source

Similar papers

#machine learning Preprint Sep 2026

Attention Manifolds: Steering or Blocking Language Models by Editing Learned B-Spline Surfaces

In standard transformer attention, a source token sends the same value vector to every receiver. The query determines \emph{how much} to attend but not \emph{what} to extract. This work introduces \textbf{attention manifolds}: learned 2D B-spline surfaces $S_d(q_d, k_d)$ that modulate each value dimension based on the...

N. Mysore · 0 citations
Preprint Aug 2026

When does fusing hand-crafted knowledge with learned representations pay? A cost-normalized benchmark of stacking, substitution, and interference

This work benchmarked one fixed hand-crafted knowledge source, a pinned bank of Gabor targets injected only during training at $\sim$2\% overhead, against data-driven alternatives under one frozen recipe with fixed subsets.

Ahmad AlMughrabi, Albert Clop, Benjamin Busam et al. · 0 citations
#small language model Preprint Sep 2026

ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

ProCAP is proposed, a probabilistic cross-attentive prompt learning framework that improves cross-modal interaction and training stability without updating any CLIP weights: it learns both visual and textual prompt tokens and links them through stacked bidirectional multi-head cross-attention so the two branches refine...

Hiwa Azeez Abbas, Fatemeh Daneshfar, M. Abdar · 0 citations
Open access Oct 2026

Low-cost fine-tuning of SmolVLA with smoothness regularization and chunk-aware auxiliary supervision for language-conditioned manipulation

Vision-Language-Action (VLA) models offer a unified interface for language-conditioned manipulation. This paper reports a negative result: a low-cost post-training recipe for SmolVLA with LoRA, a temporal smoothness regularizer, and chunk-aware auxiliary supervision from action-variation pseudo labels does not improv...

Chen He · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.