Skip to content
Open access

XSentiFusionNet: An Explainable Cross-Modal Attention Framework For Audio-Visual Sentiment Analysis Using Hybrid Deep Learning

Jul 2026 · Journal of Intelligent Decision Making and Information Science · Vol 3, pp. 1199-1223 · 0 citations · 30 references

TL;DR

XSentiFusionNet is an end-to-end explainable multimodal framework for audio-visual sentiment analysis that incorporates cross-modal transformer-attention, reliability-aware adaptive fusion, and XAI and showed higher robustness to noise and generalization ability on different multimodal datasets.

Abstract

The exponential growth of audio-visual content created through social media, communication platforms as well as human-computer interaction has led to a need for effective multimodal sentiment analysis. Most of the multimodal frameworks have limitations in cross-modal interaction modeling, fusion strategy adaptability, interpretability, and robustness to noise. This paper introduces XSentiFusionNet, an end-to-end explainable multimodal framework for audio-visual sentiment analysis that incorporates cross-modal transformer-attention, reliability-aware adaptive fusion, and XAI. Our approach uses CNNs, BiLSTMs, and ViTs to capture emotional features while adaptively combining acoustic and visual modalities based on their reliability levels. Model transparency is achieved by integrating SHAP, LIME, Grad-CAM, and attention mechanisms. Comprehensive experiments were performed on CMU-MOSEI, MELD, and RAVDESS benchmark datasets using comparative analysis, ablation studies, cross-dataset generalization experiments, confusion matrix analysis, explainability, and robustness to noise experiments. Our framework achieved an accuracy of 94.82%, F1-score of 94.11% and a ROC-AUC score of 96.04% on the benchmark datasets performing better than existing approaches such as CNN-LSTM, Transformer Fusion, and Multimodal BERT. Additionally, the framework showed higher robustness to noise and generalization ability on different multimodal datasets while the XAI module demonstrated interpretability by highlighting key speech and facial features used for predictions. These results show that XSentiFusionNet is a reliable and efficient framework for audio-visual sentiment analysis and can be used in real-world multimodal processing and affective computing scenarios.

Read PDF

Similar papers

Open access Jul 2026

MULTI-MODAL TRANSFORMER ARCHITECTURE WITH CROSS-ATTENTION FUSION FOR ROBUST AUDIO-VISUAL SENTIMENT ANALYSIS

The proposed architecture provides an effective and interpretable framework for robust multimodal sentiment analysis and offers a promising foundation for real-world affective computing applications.

B. Ankayarkanni, D. Usha Nandini, P. Sangeetha et al. · 0 citations
Aug 2026

HAFT: Hierarchical Audio-Enhanced Fusion Transformer for Efficient Multimodal Sentiment Analysis

HAFT routes audio–visual interaction through bottleneck tokens with adaptive depth-wise gating before integrating text, jointly reducing attention cost and counteracting modality bias; the Cascaded Audio Feature Enhancement (CAFE) framework strengthens prosodic representations via multi-scale time–frequency extraction; and grouped projections, decoupled positional attention, and layer-wise parameter sharing compress the remaining overhead.

Qing Dong, Ting Lu, Xiujin Shi et al. · 0 citations
Open access Aug 2026

Bridging Visual and Textual Cues: Cross-Attention Fusion for Multi-Modal Sentiment Analysis

Uni-modal studies on sentiment analysis (SA) has made significant strides, emotions in the real world, which are typically multi-modal, encompassing audio, images, video and other contents more than text. The several modes contribute to mutual development. The accuracy of sentiment evaluation will be enhanced even more if it is possible to mine the connections between different modalities. This work presents a hybrid Deep Auto-Encoding with Convolutional Neural Networks (DAE+CNN), a multi-modal technique based on cross-attention based multi-modal fusion model for text and visuals. Text features are first extracted using the pre-training model; text context features are then extracted CNN; image characteristics are extracted using the network model; and specific emotion-related areas in images are extracted using DAE+CNN. The retrieved characteristics of the text and image are then fused using multi-modal cross-attention, and the output is classified to ascertain the emotional polarity. The model presented in this research outperforms the baseline model in the comparison study of the Fashion dataset and Deep Fashion datasets with accuracy of 86.5 % and 85.5% and F1-scores of 75.3% and 76.7%, respectively,. Furthermore, authors carried out ablation studies, which verified that multi-modal fusion sentiment evaluation outperforms single-modal sentiment analysis.

M. Yuvaraja, Dr. C. Kumuthini · 0 citations
Aug 2026

MagXCL: enhanced multimodal adaptation gate and cross-modal contrastive learning for multimodal sentiment analysis

MagXCL, a unified framework designed to improve multimodal integration through more effective interaction between verbal and non-verbal modalities, is proposed, demonstrating the effectiveness of combining AMag with CrossCL to produce more accurate and robust multimodal sentiment predictions.

Duc-Duy Duong, Cam-Van Thi Nguyen, Duc-Trong Le · 0 citations
Open access Aug 2026

Learning Adaptive Cross-Modal Interactions for Multimodal Sentiment Analysis

A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.

Chuhan Cheng, Hangcheng Wu, Junqiao Wang et al. · 0 citations