NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures
Though absolute performance remains low, finetuning on NormViz-Train improves pair accuracy relatively by up to 125% and 36% Qwen3-VL 4B and 8B respectively, showing a path forward to teach models to connect visual perception to cultural significance.
Akhila Yerukola, Fabrice Y. Harel-Canada, Simran Khanuja et al.
· 0 citations