Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Jul 2026

Adaptive prompt-guided local cross-modal alignment for zero-shot vision recognition

Pre-trained vision-language models (VLMs) like CLIP have achieved remarkable success in zero-shot visual recognition. While recent advancements leverage Large Language Models (LLMs) to generate fine-grained category descriptions in order to enhance CLIP-based models, they often suffer from significant spatial granularity mismatch (fine-grained category descriptions vs. the global image) and rely heavily on labor-intensive handcrafted prompt templates. To address these challenges, we propose the Adaptive Prompt-guided Local Cross-modal Alignment (AP-LCA) approach. Our approach introduces a novel local cross-modal alignment strategy that utilizes image cropping and similarity-based semantic contribution assessment to precisely map fine-grained descriptions to relevant local image regions. Additionally, we incorporate an in-context learning (ICL) mechanism to automate the generation of task-adaptive prompts, seamlessly evolving from manual templates to context-aware descriptions that capture diverse visual concepts. Extensive experiments across eight benchmark datasets demonstrate the clear superiority of our proposed method.

Siying Wu, Song Wu · 0 citations