TriGuard-Lite: Parameter-Efficient Harmful Meme Recognition with Harm-aware Local Regions and Gated Multimodal Fusion
Abstract
Harmful memes distribute risk signals across global scene context, localized visual cues, and embedded text, resisting unimodal detection. We propose TRIGUARD-LITE, a parameter-efficient framework that integrates three complementary views—global image semantics, five-region local crops, and embedded text features—through a learned adaptive gated fusion mechanism with consistency regularization. With frozen CLIP ViTB/32 [2] and 1.6M trainable parameters, TriGuard-Lite achieves 0.623 AUROC on a stratified 1,021-sample subset of Hateful Memes, comparing favorably with CLIP linear probe (0.576), MMBT (0.602), and ViLBERT (0.608). Ablation shows cumulative gains from local regions (+0.011), gated fusion (+0.007), and consistency regularization (+0.009). Five-crop attention approaches 3×3 grid and saliency-guided alternatives with fewer views. The gating mechanism yields modality-routing weights (text 53.5%, local 24.7%, global 21.8%); gate-deletion Spearman correlations (ρtext=0.62, ρlocal=0.41) indicate partial faithfulness. Robustness evaluation shows a 37.9% reduction in average perturbation degradation. TriGuard-Lite is training-parameter-efficient, although its multi-view design incurs higher inference cost than single-view baselines.