Jun 2026· 2026 IEEE 2nd International Conference on Electronics, Energy Systems and Power Engineering (EESPE)· pp. 345-354· 0 citations· 35 references
Abstract
Multimodal Sentiment Analysis (MSA) aims to predict human sentiment by jointly modeling complementary information from textual, acoustic, and visual modalities. However, effectively exploiting heterogeneous multimodal features remains challenging due to semantic inconsistency, temporal misalignment, and noisy modality-specific representations. To address these issues, this paper proposes a Cross-Attention-based multimodal sentiment analysis framework that explicitly models inter-modal interactions through multi-directional cross-modal attention. Specifically, modality-specific features are first projected into a unified latent space via lightweight modality encoders, after which bidirectional cross-attention is employed to capture complementary dependencies among visual, acoustic, and textual modalities. To further enhance representation learning, multi-head attention and positional encoding mechanisms are incorporated to improve cross-modal interaction modeling and temporal structure awareness. Extensive experiments on the CMU-MOSI benchmark demonstrate that the proposed framework consistently outperforms conventional fusion baselines, including concatenation-based, additive, and self-attention-based fusion strategies. Comprehensive ablation studies further verify the effectiveness of each architectural component, while qualitative attention visualization confirms that the model learns interpretable attention patterns and focuses on semantically informative regions during multimodal fusion. These results indicate that the proposed method provides an effective and interpretable solution for multimodal sentiment analysis and offers useful insights for future attention-based multimodal fusion research.
A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.
Chuhan Cheng, Hangcheng Wu, Junqiao Wang et al.· International Conference on...· 0 citations
MagXCL, a unified framework designed to improve multimodal integration through more effective interaction between verbal and non-verbal modalities, is proposed, demonstrating the effectiveness of combining AMag with CrossCL to produce more accurate and robust multimodal sentiment predictions.
A Modality Dropout strategy is first introduced at the input stage to alleviate over-reliance on a sin-gle modality and improve robustness and the proposed Hierarchical Global-Local Interaction and Refinement framework for Multimodal Sentiment Analysis (HGLIR) is proposed.
yuanyuan zhou· Poster Volume 0007 The 2026...· 0 citations
Less obvious expressions like sarcasm and implicit sentiment are handled more effectively in this work, improving interpretation in multimodal sentiment analysis of social media data.
Prashant Adakane, Amit Gaikwad· international journal of eng...· 0 citations
A Text-Guided Hyper-modality Interaction Network (TGHIN) for multimodal sentiment analysis with differentiated feature encoding strategies for each modality and a Joint-Specific Fusion (JSF) module that enables the hyper-modality representation to refocus on the core information of each modality.
Kunxia Wang, RenLei Ding, YiHan Ge et al.· Signal, Image and Video Proc...· 0 citations
The proposed architecture provides an effective and interpretable framework for robust multimodal sentiment analysis and offers a promising foundation for real-world affective computing applications.
B. Ankayarkanni, D. Usha Nandini, P. Sangeetha et al.· International journal of com...· 0 citations