Skip to content
Open access

SAMS-M: Explicit sentiment-guided alignment and multi-dimensional mutual supervision for multimodal sentiment analysis

2026 · Computer Science and Information Systems · 0 citations

TL;DR

This work proposes an explicit sentiment-guided alignment and multi-dimensional cross-modal mutual supervisionbased model for multimodal sentiment analysis that primarily employs a fine-grained sentiment–saliency directed alignment mechanism and introduces a tripartite strong contrastive learning strategy to mitigate distribution discrepancies between heterogeneous modalities within a shared latent space.

Abstract

Multimodal Sentiment Analysis leverages the fusion of heterogeneous data to achieve fine-grained emotional understanding, which finds extensive application in large-scale public opinion monitoring and data mining. However, existing methods face two key challenges: (1) cross-modal alignment suffers from redundancy and semantic drift without explicit modeling of sentiment-critical cues, inducing spurious correlations; and (2) heterogeneous representation spaces lead to imbalanced modality contributions, particularly under weak image–text correlation or sentiment inconsistency. To address these challenges, we propose an explicit sentiment-guided alignment and multi-dimensional cross-modal mutual supervisionbased model for multimodal sentiment analysis. The model primarily employs a fine-grained sentiment–saliency directed alignment mechanism, which leverages bidirectional cross-attention to couple textual sentiment cues with visual saliency, enabling precise localization of sentiment-relevant regions. Furthermore, we introduce a tripartite strong contrastive learning strategy to mitigate distribution discrepancies between heterogeneous modalities within a shared latent space, thereby enhancing cross-modal coherence and complementarity. Finally, we design a noiserobust gating-based fusion module, which, together with text augmentation and deep supervision, facilitates effective joint optimization. Experimental results show that SAMS-M obtains the best results on MVSA-Single and MSD and remains competitive on the noisier MVSA-Multiple benchmark; thus, the evidence supports strong but dataset-dependent performance rather than uniform state-of-the-art superiority.

Read PDF

Similar papers

Open access Sep 2026

Label-Aware Entropic Distributional Contrastive Alignment for Multimodal Sentiment Analysis

Multimodal sentiment analysis integrates text, audio, and visual signals to infer affective states. However, sentiment information shared across modalities is often entangled with modality-specific variation, and existing representation learning methods do not fully exploit the relations encoded by continuous sentiment...

Meng-Yao Wang, Xiu-Yang Meng, Chun-Ling Wang · 0 citations
Aug 2026

Aspect-guided dual-branch fusion network for multimodal aspect-based sentiment analysis

An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.

Bin Song, Wenjing Liu, Zhi Liang et al. · 0 citations
#small language model Preprint Aug 2026

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.

Shanshan Lin, Yuesheng Wu, Chao Chen et al. · 0 citations
Preprint Sep 2026

Multimodal Aspect-Level Sentiment Analysis Based on Gated Noise Filtering and Emotion-Relevance Interaction

A Gated Noise-filtered Sentiment-Relevance Interaction (GNSRI) framework that employs a gated noise-filtering module to suppress sentiment-irrelevant features and enhance aspect-aware sentiment cues, and a sentiment-relevance interaction module to capture consistent and conflicting cross-modal signals at micro and macr...

Chen Huang, Liang-Wei Guo, Ya-Min Li et al. · 0 citations
Open access 2026

A Bidirectional Cross-Modal Attention Framework for Multimodal Sentiment Analysis

Multimodal sentiment analysis on social media data presents unique challenges due to label noise, modality conflicts, and class imbalance inherent in annotated image-text datasets. The proposed work presents CLIP-CrossFusion Net, a novel multimodal sentiment analysis framework that integrates the Contrastive Language-I...

B. Harish, C. Roopa, M. S. Kendagannaswamy et al. · 0 citations
Sep 2026

Reliability-aware disentangled adaptive network for multimodal sentiment analysis

A novel reliability-aware disentangled adaptive network that consists of three components, dynamically modulating per-modality contributions by information quality to mitigate misleading effects of unreliable modalities is proposed.

Jia-Hao Xu, Xue-Feng Zhao, Li Jia et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.