Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning
A new multimodal ICL framework is proposed that combines contrastive demonstration modeling with the self-refinement capability of MLLMs and consistently improves MLLM performance, with particularly notable gains on visual question answering (VQA).
Ming-Bo Yang, Wen-Qiang Wang, Zhaolu Kang et al.
· 1 citation