This work investigates cross-modal information flow in a video-based generative multiple-choice-like setting by applying a layer-wise causal intervention on video-text attention pathways and identifies a distinct pattern in temporal reasoning.
Abstract
Vision-Language Models are commonly evaluated through their final predictions, but understanding whether these decisions are grounded in visual evidence requires tracing how visual information contributes to language-based decisions. With this purpose in mind, we investigate cross-modal information flow in a video-based generative multiple-choice-like setting by applying a layer-wise causal intervention on video-text attention pathways. We target spatial, causal, and temporal visual reasoning. Our results show that visual information is mainly integrated while the model processes the candidate answer options, which serve as the primary textual grounding sites for the final decision. We further show that nouns play an important role as semantic anchors during multimodal enrichment, while verbs are more relevant when temporal relations are processed. Finally, we identify a distinct pattern in temporal reasoning, suggesting that VLMs struggle to reconstruct sequential information across video frames, but we remark that such fragility may also reflect linguistic biases associated with specific temporal expressions used for defining the relation between events within a scene.
SPMC, a training framework that identifies grounding-responsive heads and selectively regularizes their text-to-image attention, demonstrates that targeted guidance of sparse and implicit visual evidence pathways can directly improve multimodal reasoning.
Jia-Qi Deng, Zong-Han Wu, Zhan Heng et al.· 0 citations
This work observes that highly causal vision tokens often lie outside the target region, and extends the analysis to larger vision-language models, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations.
S. NarenKumar, T. Bhatt, Mayank Singh· 0 citations
Vision-language models (VLMs) can answer simple visual questions, but often struggle when one question requires several visual judgments. We study this gap with controlled tasks for feature binding, numerosity, spatial relations, and amodal completion, together with a Composite task that combines them. Matched counterf...
Reasoning segmentation aims to interpret implicit textual queries and enable fine-grained visual perception, which is critical for applications such as human-computer interaction and embodied agents. Existing methods typically generate explicit Chain-of-Thought (CoT) by multimodal large language models (MLLMs) before l...
Tian-Hang Guo, Yu-Lin He, Wei Chen et al.· 0 citations
irection patching is addressed with direction patching, a class-conditioned causal intervention applied across layers, token positions, and prompt formats to reframe the encoding-grounding gap as a problem of conditional transport in VLMs.
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is partic...
Xiao-Yu Zhu, Xin-Ke Deng, Suresh Taddewadikar et al.· 0 citations
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.