Skip to content
Open access

Register-augmented attention and self-calibrated fusion for robust multimodal sentiment analysis

Sep 2026 · Multimedia Systems · Vol 32 · 0 citations · 34 references

TL;DR

Results indicate that combining register-based representation stabilization with quality-aware fusion provides an effective solution for robust multimodal sentiment analysis, and RegCal-Net achieves competitive performance across multiple evaluation metrics, particularly improving fine-grained sentiment classification and reducing prediction error.

Abstract

Multimodal sentiment analysis (MSA) aims to understand human emotions by integrating heterogeneous modalities such as language and acoustic signals. Despite recent progress, two key challenges remain: (1) inter-modal inconsistency, where different modalities may convey conflicting sentiment cues, and (2) intra-modal feature ambiguity caused by noise and subtle emotional variations. To address these issues, we propose RegCal-Net, a register-augmented and self-calibrated framework for bimodal sentiment analysis. First, we introduce a Register-Augmented Self-Attention (RASA) mechanism that appends learnable register tokens along the sequence dimension to provide auxiliary global anchors for each modality. Second, we design a Self-Calibrated Fusion (SCF) module that dynamically evaluates fused feature reliability through an auxiliary score-guided gating strategy, enabling adaptive suppression of unreliable signals during multimodal integration. Extensive experiments on two widely used benchmarks, CMU-MOSI and CMU-MOSEI, together with CMU-MOSI encoder-controlled baselines and a supplementary video-subset validation, demonstrate that RegCal-Net achieves competitive performance across multiple evaluation metrics, particularly improving fine-grained sentiment classification and reducing prediction error. These results indicate that combining register-based representation stabilization with quality-aware fusion provides an effective solution for robust multimodal sentiment analysis.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis

This work introduces an adaptive variational information bottleneck to model modality-wise uncertainty and perform quality-aware information compression, thereby suppressing redundant noise in unreliable modalities and designs a reliability-aware cross-sample enhancement strategy that retrieves high-confidence, semanti...

Meng-Hua Jiang, Hao-Kai Gao, Xian-Gui Kang et al. · 0 citations
Open access Aug 2026

Learning Adaptive Cross-Modal Interactions for Multimodal Sentiment Analysis

A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.

Chuhan Cheng, Hangcheng Wu, Jun-Qiao Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MVFA: A Multi-View Text-Guided Multimodal Fusion LLM Adapter for Sentiment Analysis and Emotion Recognition

The multi-view text-guided multimodal fusion adapter (MVFA) is proposed, a parameter-efficient framework that augments frozen LLMs with strong multimodal reasoning capability and achieves state-of-the-art performance on key metrics while updating only a small fraction of parameters.

Peng-Fei Shao, Ji-Sheng Dang, Jia-Wen Fang et al. · 0 citations
Sep 2026

Reliability-aware disentangled adaptive network for multimodal sentiment analysis

A novel reliability-aware disentangled adaptive network that consists of three components, dynamically modulating per-modality contributions by information quality to mitigate misleading effects of unreliable modalities is proposed.

Jia-Hao Xu, Xue-Feng Zhao, Li Jia et al. · 0 citations
Open access Sep 2026

A Quantum-Inspired Residual Self-Attention Network for Multimodal Sentiment Analysis

Multimodal sentiment analysis is challenging because textual, visual, and acoustic evidence is heterogeneous and oft en weakly aligned. Here, we present QRSAN, a quantum-inspired residual self-attention network that integrates an LSTM text encoder, modality-specific multilayer perceptrons for visual and acoustic inputs...

Yu-Peng Liu, Xian-Jie Feng, Ye-Wang Zhong · 0 citations
#small language model Preprint Aug 2026

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.

Shanshan Lin, Yuesheng Wu, Chao Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.