Skip to content
Book Open access

Vidyashilp_Tech_Explorers@CC-MMD 2026: Cross-Cultural Misogynistic Meme Detection using Multimodal Large Language Models

Oct 2026 · Proceedings of the 28th International Conference On Multimodal Interaction · 1 citation · 21 references

Abstract

Automated detection of misogynistic content in memes presents a unique challenge at the intersection of multimodal reasoning, multilingual understanding and cultural subjectivity. We present our submission to the CC-MMD 2026 Cross-Cultural Multimodal Misogyny Detection Grand Challenge, which requires simultaneous prediction of misogyny labels from Indian, Chinese, and Western cultural perspectives across English, Chinese, Tamil, and Malayalam memes. We propose a Cross Cultural TriModal (CCTM) architecture that fine-tunes the VLM using QLoRA [3] with three parallel culture-specific classification heads, Entity-Anchored Sociocultural Context (EASC) for cultural knowledge injection, generation supervision and Group Distributionally Robust Optimization. Our system achieves Rank 1 on Malayalam (F1: 0.9372), Rank 2 on English (F1: 0.8016), Rank 4 on Chinese (F1: 0.8553) and Rank 7 on Tamil (F1: 0.7909) on the CC-MMD 2026 leaderboard. Code is available at: https://github.com/arnoldsachith/Vidyashilp_Tech_Explorers-CC-MMD-2026.git

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.