Preprint
Aug 2026
Where To Look? : Causal Tracing of Vision Encoders in VLM
This work observes that highly causal vision tokens often lie outside the target region, and extends the analysis to larger vision-language models, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations.
Narendra Kumar, Tirth Bhatt, Mayank Singh
· 0 citations