Adaptive prompt-guided local cross-modal alignment for zero-shot vision recognition
Pre-trained vision-language models (VLMs) like CLIP have achieved remarkable success in zero-shot visual recognition. While recent advancements leverage Large Language Models (LLMs) to generate fine-grained category descriptions in order to enhance CLIP-based models, they often suffer from significant spatial granularity mismatch (fine-grained category descriptions vs. the global image) and rely heavily on labor-intensive handcrafted prompt templates. To address these challenges, we propose the Adaptive Prompt-guided Local Cross-modal Alignment (AP-LCA) approach. Our approach introduces a novel local cross-modal alignment strategy that utilizes image cropping and similarity-based semantic contribution assessment to precisely map fine-grained descriptions to relevant local image regions. Additionally, we incorporate an in-context learning (ICL) mechanism to automate the generation of task-adaptive prompts, seamlessly evolving from manual templates to context-aware descriptions that capture diverse visual concepts. Extensive experiments across eight benchmark datasets demonstrate the clear superiority of our proposed method.