Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations dire...
Yanshu Li, Jia-Qian Li, Can-Ran Xiao et al.· 0 citations
Vision foundation models are increasingly reused as frozen backbones for downstream visual recognition, making parameter-efficient adaptation a central problem. Prompt-based adaptation, including Visual Prompt Tuning (VPT), provides a lightweight way to specialize these models, but its layer-wise behavior remains poorl...
Yuqi Li, Xi Xiao, Yunbei Zhang et al.· arXiv.org· 4 citations
This work injects two complementary semantic priors into Visual prompt tuning, a cascaded scheme that integrates both priors throughout ViT adaptation, and proposes a cascaded scheme that integrates both priors throughout ViT adaptation.
Xi Xiao, Xing-Jian Li, Cheng Han et al.· Trans. Mach. Learn. Res.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.