2026· Computer Science and Information Systems· 0 citations
TL;DR
This work proposes an explicit sentiment-guided alignment and multi-dimensional cross-modal mutual supervisionbased model for multimodal sentiment analysis that primarily employs a fine-grained sentiment–saliency directed alignment mechanism and introduces a tripartite strong contrastive learning strategy to mitigate distribution discrepancies between heterogeneous modalities within a shared latent space.
Abstract
Multimodal Sentiment Analysis leverages the fusion of heterogeneous data to achieve fine-grained emotional understanding, which finds extensive application in large-scale public opinion monitoring and data mining. However, existing methods face two key challenges: (1) cross-modal alignment suffers from redundancy and semantic drift without explicit modeling of sentiment-critical cues, inducing spurious correlations; and (2) heterogeneous representation spaces lead to imbalanced modality contributions, particularly under weak image–text correlation or sentiment inconsistency. To address these challenges, we propose an explicit sentiment-guided alignment and multi-dimensional cross-modal mutual supervisionbased model for multimodal sentiment analysis. The model primarily employs a fine-grained sentiment–saliency directed alignment mechanism, which leverages bidirectional cross-attention to couple textual sentiment cues with visual saliency, enabling precise localization of sentiment-relevant regions. Furthermore, we introduce a tripartite strong contrastive learning strategy to mitigate distribution discrepancies between heterogeneous modalities within a shared latent space, thereby enhancing cross-modal coherence and complementarity. Finally, we design a noiserobust gating-based fusion module, which, together with text augmentation and deep supervision, facilitates effective joint optimization. Experimental results show that SAMS-M obtains the best results on MVSA-Single and MSD and remains competitive on the noisier MVSA-Multiple benchmark; thus, the evidence supports strong but dataset-dependent performance rather than uniform state-of-the-art superiority.
Multimodal sentiment analysis integrates text, audio, and visual signals to infer affective states. However, sentiment information shared across modalities is often entangled with modality-specific variation, and existing representation learning methods do not fully exploit the relations encoded by continuous sentiment...
An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.
Bin Song, Wenjing Liu, Zhi Liang et al.· Signal, Image and Video Proc...· 0 citations
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations
A Gated Noise-filtered Sentiment-Relevance Interaction (GNSRI) framework that employs a gated noise-filtering module to suppress sentiment-irrelevant features and enhance aspect-aware sentiment cues, and a sentiment-relevance interaction module to capture consistent and conflicting cross-modal signals at micro and macr...
Chen Huang, Liang-Wei Guo, Ya-Min Li et al.· 0 citations
Multimodal sentiment analysis on social media data presents unique challenges due to label noise, modality conflicts, and class imbalance inherent in annotated image-text datasets. The proposed work presents CLIP-CrossFusion Net, a novel multimodal sentiment analysis framework that integrates the Contrastive Language-I...
B. Harish, C. Roopa, M. S. Kendagannaswamy et al.· International Journal of Adv...· 0 citations
A novel reliability-aware disentangled adaptive network that consists of three components, dynamically modulating per-modality contributions by information quality to mitigate misleading effects of unreliable modalities is proposed.
Jia-Hao Xu, Xue-Feng Zhao, Li Jia et al.· International Journal of Dat...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.