Can't Find Waldo: Evaluating VLMs'Sensitivity to Image Resolution and Detail Level
Visual Language Models (VLMs) have achieved remarkable success across diverse tasks, yet they struggle with high-resolution inputs where critical information resides in small regions or detailed, cluttered scenes. While several approaches address this limitation, a systematic understanding of why models fail at high re...