Results indicate that combining register-based representation stabilization with quality-aware fusion provides an effective solution for robust multimodal sentiment analysis, and RegCal-Net achieves competitive performance across multiple evaluation metrics, particularly improving fine-grained sentiment classification and reducing prediction error.
Abstract
Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating heterogeneous modalities such as language and acoustic signals. Despite recent progress, two key challenges remain: (1) inter-modal inconsistency, where different modalities may convey conflicting sentiment cues, and (2) intra-modal feature ambiguity caused by noise and subtle emotional variations. To address these issues, we propose RegCal-Net, a register-augmented and self-calibrated framework for bimodal sentiment analysis. First, we introduce a Register-Augmented Self-Attention (RASA) mechanism that appends learnable register tokens along the sequence dimension to provide auxiliary global anchors for each modality. Second, we design a Self-Calibrated Fusion (SCF) module that dynamically evaluates fused feature reliability through an auxiliary score-guided gating strategy, enabling adaptive suppression of unreliable signals during multimodal integration. Extensive experiments on two widely used benchmarks, CMU-MOSI and CMU-MOSEI, together with CMU-MOSI encoder-controlled baselines and a supplementary video-subset validation, demonstrate that RegCal-Net achieves competitive performance across multiple evaluation metrics, particularly improving fine-grained sentiment classification and reducing prediction error. These results indicate that combining register-based representation stabilization with quality-aware fusion provides an effective solution for robust multimodal sentiment analysis.
This work introduces an adaptive variational information bottleneck to model modality-wise uncertainty and perform quality-aware information compression, thereby suppressing redundant noise in unreliable modalities and designs a reliability-aware cross-sample enhancement strategy that retrieves high-confidence, semanti...
Meng-Hua Jiang, Hao-Kai Gao, Xian-Gui Kang et al.· 0 citations
A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.
Chuhan Cheng, Hangcheng Wu, Jun-Qiao Wang et al.· International Conference on...· 0 citations
The multi-view text-guided multimodal fusion adapter (MVFA) is proposed, a parameter-efficient framework that augments frozen LLMs with strong multimodal reasoning capability and achieves state-of-the-art performance on key metrics while updating only a small fraction of parameters.
Peng-Fei Shao, Ji-Sheng Dang, Jia-Wen Fang et al.· 0 citations
A novel reliability-aware disentangled adaptive network that consists of three components, dynamically modulating per-modality contributions by information quality to mitigate misleading effects of unreliable modalities is proposed.
Jia-Hao Xu, Xue-Feng Zhao, Li Jia et al.· International Journal of Dat...· 0 citations
Multimodal sentiment analysis is challenging because textual, visual, and acoustic evidence is heterogeneous and oft en weakly aligned. Here, we present QRSAN, a quantum-inspired residual self-attention network that integrates an LSTM text encoder, modality-specific multilayer perceptrons for visual and acoustic inputs...
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.