Skip to content

NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures

Sep 2026 · 0 citations · 70 references
Computer Science

TL;DR

Though absolute performance remains low, finetuning on NormViz-Train improves pair accuracy relatively by up to 125% and 36% Qwen3-VL 4B and 8B respectively, showing a path forward to teach models to connect visual perception to cultural significance.

Abstract

AI systems are used worldwide, but they struggle to serve the needs of culturally diverse populations. Prior work on cultural understanding evaluates AI systems on text-only settings or on visual artifact recognition (e.g. foods, clothing). The ability to reason about visually observable behaviors through local social norms, which we call visual norm understanding, remains unexamined. We introduce NormViz-Bench, a high quality, human-validated benchmark of 3,268 contrastive image pairs (6,536 images) spanning 16 countries. Each pair varies only in the culturally relevant behavior (e.g., objects, attributes, spatial relations, and actions) that alters how each image is interpreted. Each image is labeled as conforming to, violating, or irrelevant to local social norms, and pair-level evaluation requires both images to be correctly classified, thereby preventing reliance on superficial visual shortcuts. Even the strongest VLMs, Gemini 3.0 Flash and Qwen2.5 VL 7B, succeed on only 26.6% and 21.6% of pairs, struggling most with identifying violating and culturally benign visual behaviors. Towards bridging this, we introduce NormViz-Train, a training dataset of 64k images paired with explanations. Though absolute performance remains low (<30%), finetuning on NormViz-Train improves pair accuracy relatively by up to 125% and 36% Qwen3-VL 4B and 8B respectively, showing a path forward to teach models to connect visual perception to cultural significance. Together, NormViz-Bench and NormViz-Train establish visual norm understanding as a challenging and consequential frontier for multimodal AI.

View source

Similar papers

Preprint Aug 2026

Human-AI Perceptual Alignment by Playing Hues and Cues

Evaluating the perceptual alignment between Contrastive Vision-Language Models (CVLMs) and humans is typically constrained by traditional benchmarks that overlook fine-grained semantic and cultural nuances. In this work, we propose a novel evaluation framework that leverages the gamified, discrete color space of the bo...

Nuria Alabau-Bosque, Jorge Vila-Tomás, Paula Daudén-Oliver et al. · 0 citations
Preprint Aug 2026

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

PoVisLE is introduced, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context.

Anna Kołos, Grzegorz Statkiewicz, Karolina Seweryn et al. · 1 citation
Preprint Aug 2026

HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

HUG-VIS, a unified benchmark for Human-centered Understanding and Generation in Visual Intelligence, contains 8,400 seated half-body videos of 30 professional actors, each performing the same 280 emotion-action-prompt assignments under a controlled Mandarin studio protocol, with synchronized video, audio, text, and alp...

Fei Ma, Ze-Bang Cheng, Ming-Hui Li et al. · 0 citations
Open access Aug 2026

Inquiry That Holds Up in a World of AI Deepfakes and Synthetic Text: A Verification-Ready Routine for K-12 Social Studies

AI-generated images, audio, video, and text are now part of the information environment students encounter every day. Students increasingly face claims that look credible while remaining misleading, incomplete, or difficult to trace to a reliable source (Chesney & Citron, 2019; Vaccari & Chadwick, 2020). This article a...

Steven Grubaugh, Greg Levitt, I.Y. Lam · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.