Jul 2026· International journal of computer information systems and industrial management applications· 0 citations
TL;DR
The proposed architecture provides an effective and interpretable framework for robust multimodal sentiment analysis and offers a promising foundation for real-world affective computing applications.
Abstract
Multimodal sentiment analysis (MSA) has gained significant attention due to its ability to integrate heterogeneous information from audio, visual, and textual modalities. However, existing transformer-based fusion methods often suffer from reduced robustness when one or more modalities are corrupted or partially unavailable. This paper presents a Multi-Modal Transformer Architecture with Cross-Attention Fusion (MMT-CAF) for robust audio-visual sentiment analysis. The proposed framework combines modality-specific transformer encoders, bidirectional cross-attention, and a reliability-aware fusion mechanism that dynamically adjusts the contribution of each modality according to its estimated reliability. The framework was evaluated on the CMU-MOSI and CMU-MOSEI benchmark datasets and compared with representative transformer-based methods, including Adaptive Modality Weighting, RAFT, and CITN-DAF. Experimental results demonstrate that MMT-CAF achieves superior sentiment classification performance while maintaining higher robustness under noisy audio, visual occlusion, and missing-modality scenarios. Ablation studies further confirm the effectiveness of the proposed cross-attention and reliability-aware fusion modules in improving multimodal representation learning. The proposed architecture provides an effective and interpretable framework for robust multimodal sentiment analysis and offers a promising foundation for real-world affective computing applications.
HAFT routes audio–visual interaction through bottleneck tokens with adaptive depth-wise gating before integrating text, jointly reducing attention cost and counteracting modality bias; the Cascaded Audio Feature Enhancement (CAFE) framework strengthens prosodic representations via multi-scale time–frequency extraction; and grouped projections, decoupled positional attention, and layer-wise parameter sharing compress the remaining overhead.
Qing Dong, Ting Lu, Xiujin Shi et al.· International journal of sof...· 0 citations
XSentiFusionNet is an end-to-end explainable multimodal framework for audio-visual sentiment analysis that incorporates cross-modal transformer-attention, reliability-aware adaptive fusion, and XAI and showed higher robustness to noise and generalization ability on different multimodal datasets.
M. Kidwai, C. Author, Dr. Faiyaz Ahmad· Journal of Intelligent Decis...· 0 citations
A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.
Chuhan Cheng, Hangcheng Wu, Junqiao Wang et al.· International Conference on...· 0 citations
Uni-modal studies on sentiment analysis (SA) has made significant strides, emotions in the real world, which are typically multi-modal, encompassing audio, images, video and other contents more than text. The several modes contribute to mutual development. The accuracy of sentiment evaluation will be enhanced even more if it is possible to mine the connections between different modalities. This work presents a hybrid Deep Auto-Encoding with Convolutional Neural Networks (DAE+CNN), a multi-modal technique based on cross-attention based multi-modal fusion model for text and visuals. Text features are first extracted using the pre-training model; text context features are then extracted CNN; image characteristics are extracted using the network model; and specific emotion-related areas in images are extracted using DAE+CNN. The retrieved characteristics of the text and image are then fused using multi-modal cross-attention, and the output is classified to ascertain the emotional polarity. The model presented in this research outperforms the baseline model in the comparison study of the Fashion dataset and Deep Fashion datasets with accuracy of 86.5 % and 85.5% and F1-scores of 75.3% and 76.7%, respectively,. Furthermore, authors carried out ablation studies, which verified that multi-modal fusion sentiment evaluation outperforms single-modal sentiment analysis.
M. Yuvaraja, Dr. C. Kumuthini· Journal of Intelligent Decis...· 0 citations
The Gate-Controlled Information Bottleneck Cross-Modal Attention Network (GICA), a two-stage hierarchical framework that jointly optimizes adaptive modal compression and cross-modal information interaction, is offered, addressing the issues of over-compression and under-compression present in current information bottleneck techniques.
Yu-Hao Qiang, Junfeng Shen· Pattern Analysis and Applica...· 0 citations