BanglaMemeX is introduced, a culturally grounded multimodal benchmark comprising 3,000 Bangla memes annotated with multi-dimensional labels and human-written explanations that explicitly describe textual and visual metaphors.
Abstract
Vision Language Models have achieved strong performance on multimodal benchmarks, yet their ability to reason about culturally grounded and metaphor-rich content remains insufficiently studied. Internet memes present a challenging setting where meaning emerges from implicit interactions between image, overlaid text, sarcasm, and shared socio-cultural knowledge rather than literal visual recognition. This challenge is amplified in low-resource languages such as Bangla, where code-mixing, stylized scripts, and culturally specific symbolism introduce substantial distribution shift. In this work, we introduce BanglaMemeX, a culturally grounded multimodal benchmark comprising 3,000 Bangla memes annotated with multi-dimensional labels (humor, sarcasm, offensiveness, motivational intent, and overall sentiment) and human-written explanations that explicitly describe textual and visual metaphors. We systematically evaluate modern VLMs on both classification and explanation generation, revealing that current models struggle to interpret implicit cultural cues despite reasonable surface-level accuracy. Our results highlight the need for culturally-aware multimodal systems capable of grounded reasoning under linguistic and cultural distribution shift.
This work introduces MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context note and three human-written explanations, along with a supplementary set of 54 Bengali regional dialect memes.
Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir et al.· 0 citations
PoVisLE is introduced, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context.
Anna Kołos, Grzegorz Statkiewicz, Karolina Seweryn et al.· 1 citation
An overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation is presented, covering spoken visual question answering and image-grounded hallucination detection in English and Modern Standard Arabic, and CRAI-Bench, evaluating the cultural accuracy of text-to-image generation.
CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context.
Bo Zeng, Lin-Feng Gao, Pei-Qing Lin et al.· 0 citations
Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored. Prior work has focused largely on modern language varieties and poetry, while classical prose traditions such as ma...
AbdulRahman A. Morsy, Aya Zirikly Department of Computer Science, School Of Electrical Engineering et al.· 0 citations
PUMA (Polish Unified Multimodal Assessment) is proposed, a novel benchmark of 900 hand-crafted tasks designed to probe the limits of multimodal models in the Polish cultural and linguistic context and open-source the evaluation framework to advance localized multimodal AI research.
Slawomir Dadas, Michał Perełkiewicz, Rafal Poswiata et al.· 0 citations
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.