Skip to content
Preprint

Who Drives the Probability Game of VLMs? A Temporal Causal Drive Evaluation Framework

Sep 2026 · 0 citations · 46 references
Computer Science

TL;DR

Results show that causal-drive trajectories provide complementary source-level diagnostics for multimodal generation and show a consistent transition from stronger early question and visual guidance toward increasing reliance on generated prefixes.

Abstract

Vision-language models (VLMs) are increasingly evaluated on complex image and video understanding tasks, yet conventional metrics primarily assess final-answer quality and reveal little about how different information sources shape the generation process. We propose a causal and temporal evaluation framework that traces the evolving roles of visual input, question text, and generated prefixes during autoregressive decoding. Grounded in a Structural Causal Model, we use interventions and backdoor adjustment to derive three step-indexed causal-drive metrics---Visual Causal Drive (VCD), Question Causal Drive (QCD), and Prefix Causal Drive (PCD)---for characterizing source-specific generation patterns without requiring reference answers. Experiments on Qwen3-VL-8B-Instruct across MAVIS, LLaVA-Video-178K, and MiraData, together with cross-model validation on InternVL2-8B, reveal a consistent transition from stronger early question and visual guidance toward increasing reliance on generated prefixes. Randomized-intervention validation shows that QCD and PCD reduce recovery error over observational PMI baselines by 34.8\% and 47.1\%, respectively. On VLMBias, the prefix--visual imbalance score achieves 0.767 AUROC and 0.873 AUPRC for distinguishing prior-driven from visually grounded generations. These results show that causal-drive trajectories provide complementary source-level diagnostics for multimodal generation.

View source

Similar papers

#machine learning Preprint Aug 2026

Do VLMs Share Safety Neurons Across Modalities?

A causal, neuron-level analysis of safety mechanisms in 10 VLMs, a two-stage detection pipeline with iterative ablation that accounts for self-repair, and two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval, which decouple visual and textual safety signals are introduced.

Jia-Xuan Li, Jia-Hao Zhang, D. Vo et al. · 0 citations
Preprint Aug 2026

Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding

Clue-OPSD, a clue-privileged on-policy self-distillation framework for long-video understanding that uses clue intervals as privileged supervision without relying on ground-truth answer labels, while requiring no clue annotations or additional modules at inference time is introduced.

Kaishen Wang, Dong-Di Zhao, Yijun Liang et al. · 4 citations
#computer vision Preprint Aug 2026

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

It is indicated that current VLMs can identify anomalies within individual frames but struggle to integrate information across frames to reason about temporal consistency, and TimeCatch provides a controlled benchmark for evaluating temporal grounding in vision-language models.

Marek Hradil, Danae Sánchez Villegas · 0 citations
Preprint Sep 2026

CCRV-Bench: Constraint-Based Evaluation of Causal Reasoning in Vision-Language Models

Vision-language models (VLMs) have demonstrated excellent performance in visual tasks, but their visual causal reasoning capabilities still lack reliable evaluation. Existing evaluations struggle to distinguish whether a model is performing causal reasoning based on visual evidence or relying on statistical correlation...

Lin-Yuan Gao, Yuan Wu, Yi Chang · 0 citations
Preprint Aug 2026

How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures

SciFigBench is introduced, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty and proposes the Admittance-Resistance-Inductance (A-R-I) framework to evaluate whether models acknowledge insufficient evidence, resist mi...

Paul Osemudiame Oamen, Owusu-Banahene Osei, Ananya Mukherjee et al. · 0 citations
#computer vision Preprint Aug 2026

STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models

STRAND is introduced, a benchmark of human-verified object-centric facts that evaluates intermediate reasoning by decomposing queries into sub-questions, distinguishing genuine temporal understanding from coincidental correctness and an object-centric framework that explicitly constructs and reasons over structured obj...

T. Nguyen, Tri Cao, Khoi M. Le et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.