Skip to content
Open access

Learning Adaptive Cross-Modal Interactions for Multimodal Sentiment Analysis

Aug 2026 · International Conference on Data Technologies and Applications · 0 citations · 39 references

TL;DR

A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.

Abstract

Multimodal sentiment analysis aims to integrate heterogeneous textual, visual, and acoustic information for effective emotion understanding. However, existing methods often suffer from insufficient cross-modal interaction modeling, limited adaptability in multimodal fusion, and inadequate suppression of modality-specific noise under complex conversational scenarios. To address these challenges, this paper proposes a framework for learning adaptive cross-modal interactions for multimodal sentiment analysis. The proposed framework consists of three stages: modality-aware preprocessing, heterogeneous representation learning, and adaptive multimodal fusion. First, a unified preprocessing strategy is designed to improve cross-modal consistency through textual normalization, speaker-aware visual alignment, and utterance-level acoustic representation enhancement. Second, modality-specific encoders are constructed to capture complementary semantic, spatial, and utterance-level acoustic characteristics from textual, visual, and acoustic modalities, respectively. Third, an adaptive fusion framework is introduced to explicitly model cross-modal interactions, dynamically estimate the importance of different modality combinations, and further calibrate discriminative feature channels through channel attention. By jointly performing modality-level interaction learning and channel-wise feature refinement, the proposed framework effectively enhances multimodal representation capability for sentiment classification. Extensive experiments conducted on the CMU-MOSI and MELD benchmark datasets demonstrate that our framework consistently outperforms previous methods. In particular, the proposed model achieves 90.27% accuracy and 90.26% F1-score on CMU-MOSI, together with 66.57% accuracy and 66.21% F1-score on MELD. Additional ablation studies and qualitative analyses further validate the effectiveness of the proposed preprocessing strategy, modality-specific representation learning, and adaptive fusion mechanism.

Read PDF

Similar papers

Aug 2026

TGHIN: text-guided hyper-modality interaction network for multimodal sentiment analysis

A Text-Guided Hyper-modality Interaction Network (TGHIN) for multimodal sentiment analysis with differentiated feature encoding strategies for each modality and a Joint-Specific Fusion (JSF) module that enables the hyper-modality representation to refocus on the core information of each modality.

Kunxia Wang, RenLei Ding, YiHan Ge et al. · 0 citations
Aug 2026

MagXCL: enhanced multimodal adaptation gate and cross-modal contrastive learning for multimodal sentiment analysis

MagXCL, a unified framework designed to improve multimodal integration through more effective interaction between verbal and non-verbal modalities, is proposed, demonstrating the effectiveness of combining AMag with CrossCL to produce more accurate and robust multimodal sentiment predictions.

Duc-Duy Duong, Cam-Van Thi Nguyen, Duc-Trong Le · 0 citations
Conference 2026

Hierarchical Global-Local Interaction and Refinement for Multimodal Sentiment Analysis

A Modality Dropout strategy is first introduced at the input stage to alleviate over-reliance on a sin-gle modality and improve robustness and the proposed Hierarchical Global-Local Interaction and Refinement framework for Multimodal Sentiment Analysis (HGLIR) is proposed.

yuanyuan zhou · 0 citations
Aug 2026

Discriminative semantic learning for incomplete multimodal sentiment analysis

The framework first introduces learnable sentiment prototypes as semantic anchors to provide explicit sentiment-discriminative guidance for feature completion, and a gradient decoupling strategy is designed to separate the optimization paths of unimodal and multimodal objectives, preventing fusion gradients from interfering with unimodal encoders, thereby synergistically enhancing both discriminative representation learning and multimodal fusion.

Shan Tao, Haipeng Chen, Yu Liu et al. · 0 citations
Conference Jul 2026

Multimodal AI Framework for Music Emotion and Media Sentiment Analysis using Deep Learning

With the evolution of multimedia platforms and online streaming services, the need for intelligent systems have increased to learn from heterogeneous data sources to understand human emotions and sentiments. Conventional unimodal paradigm for music emotion recognition and sentiment analysis, which relies solely on the use of single mode data (usually audio modality), has showed limited success in the task. Towards overcoming these challenges, this paper introduces an innovative cross-modal transformer based multimodal deep learning framework for combined music emotion and media sentiment analysis using synchronized multimodal representations from the CMU-MOSEI dataset. It proposes a framework which deploys Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) models to extract music-inspired acoustic emotional features, Bidirectional Encoder Representations from Transformers (BERT) for textual sentiment representation learning, and Vision Transformer (ViT)-based feature extraction for visual emotional understanding respectively. The cross-modal transformer attention fuses heterogeneous modality representations and enhances the contextual interaction learning. Experimental evaluation shows that the proposed framework significantly outperforms traditional unimodal and multimodal approaches in terms of accuracy, precision, recall and F1-score. The proposed system is a solid and scalable solution for next generation affective multimedia analytics, intelligent recommendation systems and emotion-aware digital media applications.

Madhur Thapliyal, Anuja Rohilla, Ashish Kulshrestha et al. · 0 citations
Open access Jul 2026

MULTI-MODAL TRANSFORMER ARCHITECTURE WITH CROSS-ATTENTION FUSION FOR ROBUST AUDIO-VISUAL SENTIMENT ANALYSIS

The proposed architecture provides an effective and interpretable framework for robust multimodal sentiment analysis and offers a promising foundation for real-world affective computing applications.

B. Ankayarkanni, D. Usha Nandini, P. Sangeetha et al. · 0 citations