Skip to content

From Vision to Language: Investigating Causal Information Flow in Multimodal Decision-Making

Sep 2026 · 0 citations · 36 references
Computer Science

TL;DR

This work investigates cross-modal information flow in a video-based generative multiple-choice-like setting by applying a layer-wise causal intervention on video-text attention pathways and identifies a distinct pattern in temporal reasoning.

Abstract

Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-based decisions. With this purpose in mind, we investigate cross-modal information flow in a video-based generative multiple-choice-like setting by applying a layer-wise causal intervention on video-text attention pathways. We target spatial, causal, and temporal visual reasoning. Our results show that visual information is mainly integrated while the model processes the candidate answer options, which serve as the primary textual grounding sites for the final decision. We further show that nouns play an important role as semantic anchors during multimodal enrichment, while verbs are more relevant when temporal relations are processed. Finally, we identify a distinct pattern in temporal reasoning, suggesting that VLMs struggle to reconstruct sequential information across video frames, but we remark that such fragility may also reflect linguistic biases associated with specific temporal expressions used for defining the relation between events within a scene.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

SPMC, a training framework that identifies grounding-responsive heads and selectively regularizes their text-to-image attention, demonstrates that targeted guidance of sparse and implicit visual evidence pathways can directly improve multimodal reasoning.

Jia-Qi Deng, Zong-Han Wu, Zhan Heng et al. · 0 citations
Preprint Aug 2026

Where To Look? : Causal Tracing of Vision Encoders in VLM

This work observes that highly causal vision tokens often lie outside the target region, and extends the analysis to larger vision-language models, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations.

S. NarenKumar, T. Bhatt, Mayank Singh · 0 citations
#machine learning Preprint Sep 2026

From Perception to Integration: Revisiting the Internal Dynamics of Reasoning in Vision-Language Models

Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterf...

Rong-Yu Xu, Prayag Tiwari, Shao-Lei Zhang · 0 citations
#artificial intelligence Preprint Sep 2026

Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models

Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before l...

Tian-Hang Guo, Yu-Lin He, Wei Chen et al. · 0 citations
Preprint Aug 2026

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is partic...

Xiao-Yu Zhu, Xin-Ke Deng, Suresh Taddewadikar et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.