Skip to content
Preprint

ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language Models

Aug 2026 · 1 citation · 38 references
Computer Science

TL;DR

This work proposes ZOMP (Zeroth-Order Multimodal Prompt tuning), a query-efficient, fully forward-only method that tunes deep prompts in both the vision and text branches of a frozen CLIP model using simultaneous perturbation stochastic approximation.

Abstract

Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pass access is available, as is common for memory-constrained edge devices and proprietary model deployments. Prior BP-free, zeroth-order prompt-tuning methods avoid this requirement but often tune prompts in a single modality or optimize over a search space large enough that convergence requires thousands of forward passes, which is impractical under realistic query budgets. We propose ZOMP (Zeroth-Order Multimodal Prompt tuning), a query-efficient, fully forward-only method that tunes deep prompts in both the vision and text branches of a frozen CLIP model using simultaneous perturbation stochastic approximation. ZOMP combines three ingredients: a cross-modal low-rank reparameterization that ties the two branches through a shared factor and keeps the effective search dimensionality small, a gradient-correction momentum term that stabilizes the noisy zeroth-order estimate, and a budget-indexed rank schedule that unlocks capacity as the query budget is spent. Across 13 vision-language benchmarks under a matched 5,000-query budget, ZOMP consistently outperforms prior BP-free prompt-tuning methods in both few-shot accuracy and query efficiency, and it generalizes better across base-to-new, cross-dataset transfer, and out-of-distribution settings. Our results show that jointly exploiting multimodality and low-rank structure is an effective route to practical, query-efficient BP-free prompt tuning.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

TerMeZO: Ternary Sparse Zeroth-Order Optimization for Fine-tuning BitNet Models at the Edge

Fine-tuning anguage models (LLMs) with first-order optimizers requires a memory several times larger than that required for inference. Memory-efficient zeroth-order optimization (MeZO) sidesteps this cost by estimating gradients from forward passes only. However, for BitNet architectures, a family of LLMs with ternary...

Houssem Sifaou, Prabodh Katti, B. Rajendran et al. · 1 citation
Preprint Aug 2026

SubZero+: Memory-Efficient Adaptive Zeroth-Order LLM Fine-Tuning in Random Subspaces

Zeroth-order (ZO) optimization with SGD in random subspaces enables memory-efficient fine-tuning of large language models without backpropagation. However, high gradient estimation noise fundamentally undermines adaptive optimizers like Adam. We propose SubZero+, which achieves practical adaptive ZO optimization throug...

Zi-Ming Yu, Shu-Yao Xiao, Xingyu Zhao et al. · 0 citations
#machine learning Preprint Sep 2026

Beyond Reconstruction Loss in Post-Training Quantization: Balanced Fitting for Large Vision-Language Models

Experiments on multiple LVLMs show that the Balanced Fitting method consistently outperforms prior PTQ approaches under both weight-only and weight-activation quantization, while lower reconstruction loss does not reliably translate into better downstream performance.

Min-Chan Kang, Kyeonghye Park, Seungyeon Sa et al. · 0 citations
#small language model Preprint Sep 2026

ProCAP: Probabilistic Cross-Attentive Prompt Learning for Vision-Language Models

ProCAP is proposed, a probabilistic cross-attentive prompt learning framework that improves cross-modal interaction and training stability without updating any CLIP weights: it learns both visual and textual prompt tokens and links them through stacked bidirectional multi-head cross-attention so the two branches refine...

Hiwa Azeez Abbas, Fatemeh Daneshfar, M. Abdar · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.