Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed Illusion-Reasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on Illusion-Reasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future directions for optimisation. Our project is publicly available at https://github.com/zhaoliangjie55/EMNLP2026_Illusion.
Liang Zhao, Jiaqing Lyu, Kexin Tang et al.· 0 citations
The findings reveal that transformer-based large language models exhibit learning patterns similar to those of human learners, with a faster learning speed for simpler subtasks compared to more complex ones, which suggests that transformer-based LLMs may share cognitive processes with human learners in arithmetic.
Luyu Qiu, Jianing Li, Hwanhee Kim et al.· 0 citations