Skip to content

Beyond Surface Imitation: Contrastive Modeling for Reasoning Path Alignment in Multimodal In-Context Learning

Sep 2026 · 1 citation · 60 references
Computer Science

TL;DR

A new multimodal ICL framework is proposed that combines contrastive demonstration modeling with the self-refinement capability of MLLMs and consistently improves MLLM performance, with particularly notable gains on visual question answering (VQA).

Abstract

In-context learning (ICL) is widely used in multimodal large language models (MLLMs) and achieves strong performance across a wide range of multimodal tasks. However, existing multimodal ICL methods often rely on surface level imitation of in-context demonstrations, making it difficult for MLLMs to align their responses with the reasoning path required by the given multimodal input. This limitation becomes more pronounced in complex multimodal tasks, thereby restricting further improvements in MLLM performance. To address this issue, we propose a new multimodal ICL framework that combines contrastive demonstration modeling with the self-refinement capability of MLLMs. Specifically, our framework reformulates each demonstration by explicitly contrasting a suboptimal response with a better response under the same input, together with a reasoning path that reveals how the response should be refined. This contrastive formulation makes the reasoning path toward the desired response more explicit and guides the MLLM beyond superficial imitation. Furthermore, because effective refinement depends on the current response, we introduce a response-conditioned retrieval mechanism to select demonstrations whose reasoning paths are more relevant to the current response. In addition, we use a lightweight alignment controller to predict response quality and determine whether further refinement is needed. Experiments on three types of multimodal tasks show that the proposed framework consistently improves MLLM performance, with particularly notable gains on visual question answering (VQA).

View source

Similar papers

#computer vision Preprint Sep 2026

MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models

Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations dire...

Yanshu Li, Jia-Qian Li, Can-Ran Xiao et al. · 0 citations
#natural language process... Preprint Sep 2026

Func-R1: Incentivizing Mathematical Function Reasoning in Multimodal Large Language Models

Performing deliberate mathematical reasoning in visual contexts is a hallmark of advanced Multimodal Large Language Models (MLLMs) and requires a sophisticated synthesis of perceptual grounding and symbolic logic. However, in the realm of mathematical functions, our investigation reveals a critical modality interferenc...

Ming Yin, Xiaohai Wang, Dian Li et al. · 0 citations
Preprint Sep 2026

V-ICAL Bench: Evaluating Video In-Context Learning for Multimodal Agents in Interactive Environments

While In-Context Learning (ICL) enables models to adapt from exemplars without parameter updates, multimodal ICL remains largely underexplored, particularly regarding video demonstrations in interactive environments. For multimodal agents, learning from videos presents unique challenges: they must translate in-context...

Zi-Qian Fan, Shi-Bo Xu, Jun-Jie Li et al. · 0 citations
Preprint Sep 2026

Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context understanding with joint audio-visual generation in a shared latent framework. Its unified architecture rai...

Hao-Yu Zhao, Zi-Hao Zhao, Tian-Yuan Deng et al. · 0 citations
Preprint Sep 2026

VISTA: Internalizing Collective Visual Experience via On-Policy Distillation for Active Multimodal Agents

Active multimodal agents use visual tools to acquire task-relevant evidence while reasoning. Although reinforcement learning samples multiple interaction trajectories per input, outcome-based objectives primarily use the group to estimate scalar advantages, leaving complementary visual discoveries underused. We introdu...

Zheng Jiang, Hou-De Qian, Yi-Ming Chen et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.