Skip to content

BanglaMemeX: Advancing Cultural Metaphoric Image Interpretation in Bangla with a Multimodal Explainable Dataset

Sep 2026 · 0 citations · 26 references
Computer Science

TL;DR

BanglaMemeX is introduced, a culturally grounded multimodal benchmark comprising 3,000 Bangla memes annotated with multi-dimensional labels and human-written explanations that explicitly describe textual and visual metaphors.

Abstract

Vision Language Models have achieved strong performance on multimodal benchmarks, yet their ability to reason about culturally grounded and metaphor-rich content remains insufficiently studied. Internet memes present a challenging setting where meaning emerges from implicit interactions between image, overlaid text, sarcasm, and shared socio-cultural knowledge rather than literal visual recognition. This challenge is amplified in low-resource languages such as Bangla, where code-mixing, stylized scripts, and culturally specific symbolism introduce substantial distribution shift. In this work, we introduce BanglaMemeX, a culturally grounded multimodal benchmark comprising 3,000 Bangla memes annotated with multi-dimensional labels (humor, sarcasm, offensiveness, motivational intent, and overall sentiment) and human-written explanations that explicitly describe textual and visual metaphors. We systematically evaluate modern VLMs on both classification and explanation generation, revealing that current models struggle to interpret implicit cultural cues despite reasonable surface-level accuracy. Our results highlight the need for culturally-aware multimodal systems capable of grounded reasoning under linguistic and cultural distribution shift.

View source

Similar papers

#computer vision Preprint Sep 2026

MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

This work introduces MemeCULT-1K, a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, where each meme is paired with a cultural context note and three human-written explanations, along with a supplementary set of 54 Bengali regional dialect memes.

Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir et al. · 0 citations
Preprint Aug 2026

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

PoVisLE is introduced, a monocultural vision-language benchmark for Polish designed to evaluate culturally grounded multimodal understanding under a grounded evaluation paradigm, where language is interpreted in interaction with visual context.

Anna Kołos, Grzegorz Statkiewicz, Karolina Seweryn et al. · 1 citation
#artificial intelligence Review Aug 2026

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

An overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation is presented, covering spoken visual question answering and image-grounded hallucination detection in English and Modern Standard Arabic, and CRAI-Bench, evaluating the cultural accuracy of text-to-image generation.

Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?

Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored. Prior work has focused largely on modern language varieties and poetry, while classical prose traditions such as ma...

AbdulRahman A. Morsy, Aya Zirikly Department of Computer Science, School Of Electrical Engineering et al. · 0 citations
#small language model Preprint Aug 2026

PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

PUMA (Polish Unified Multimodal Assessment) is proposed, a novel benchmark of 900 hand-crafted tasks designed to probe the limits of multimodal models in the Polish cultural and linguistic context and open-source the evaluation framework to advance localized multimodal AI research.

Slawomir Dadas, Michał Perełkiewicz, Rafal Poswiata et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Jun 3, 2026

MIT researchers teach AI models to interpret charts

The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.