Skip to content
Conference

TriGuard-Lite: Parameter-Efficient Harmful Meme Recognition with Harm-aware Local Regions and Gated Multimodal Fusion

Aug 2026 · 2026 7th International Conference on Computer Vision and Data Mining (ICCVDM) · pp. 446-449 · 0 citations · 9 references

Abstract

Harmful memes distribute risk signals across global scene context, localized visual cues, and embedded text, resisting unimodal detection. We propose TRIGUARD-LITE, a parameter-efficient framework that integrates three complementary views—global image semantics, five-region local crops, and embedded text features—through a learned adaptive gated fusion mechanism with consistency regularization. With frozen CLIP ViTB/32 [2] and 1.6M trainable parameters, TriGuard-Lite achieves 0.623 AUROC on a stratified 1,021-sample subset of Hateful Memes, comparing favorably with CLIP linear probe (0.576), MMBT (0.602), and ViLBERT (0.608). Ablation shows cumulative gains from local regions (+0.011), gated fusion (+0.007), and consistency regularization (+0.009). Five-crop attention approaches 3×3 grid and saliency-guided alternatives with fewer views. The gating mechanism yields modality-routing weights (text 53.5%, local 24.7%, global 21.8%); gate-deletion Spearman correlations (ρtext=0.62, ρlocal=0.41) indicate partial faithfulness. Robustness evaluation shows a 37.9% reduction in average perturbation degradation. TriGuard-Lite is training-parameter-efficient, although its multi-view design incurs higher inference cost than single-view baselines.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.