A Bidirectional Cross-Modal Attention Framework for Multimodal Sentiment Analysis
Abstract
Multimodal sentiment analysis on social media data presents unique challenges due to label noise, modality conflicts, and class imbalance inherent in annotated image-text datasets. The proposed work presents CLIP-CrossFusion Net, a novel multimodal sentiment analysis framework that integrates the Contrastive Language-Image Pre-Training (CLIP) ViT-B/32 image encoder with a RoBERTa-Base text encoder through a bidirectional cross-modal attention mechanism. The proposed approach addresses three key challenges: 1) rich cross-modal feature interaction via bidirectional attention, 2) noisy label robustness through focal loss with label smoothing, and 3) class imbalance mitigation via targeted class weighting and Mixup augmentation. The study further proposes an extended variant with a colour feature branch, supervised contrastive loss, and deeper cross-modal attention layers. Using the proposed model, comprehensive experiments are conducted on both the MVSA-Multiple (19,600 image-text tweet pairs) and MVSA-Single (4,511 pairs) datasets. On MVSA-Multiple, the proposed model achieves 74.39% accuracy and a weighted F1 of 0.7497. On MVSA-Single, the proposed ensemble of three complementary model variants (base model, colour-contrastive variant, and RoBERTa-Large variant) achieved 75.66% accuracy and an F1 of 0.7517, outperforming strong baselines. These results demonstrate the effectiveness of cross-modal attention fusion combined with noise-robust training strategies for multimodal sentiment analysis.