The Alignment Illusion in Multimodal Large Language Models
Internal visual-text alignment in MLLMs is best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.
Hong-Han Wang, Yun-Tao Wang, Hui-Chao Ding
· 0 citations