Layers, Sinks, and Scaling: Adaptive Evidence Selection for Multimodal Large Language Models
Multimodal large language models (MLLMs) can answer knowledge-intensive visual questions by combining visual evidence from images with facts retrieved from external sources. However, MLLMs may overlook relevant evidence in both modalities, attending weakly to the textual sentences or visual regions needed for the corre...