Jun 2026· 2026 IEEE 2nd International Conference on Electronics, Energy Systems and Power Engineering (EESPE)· pp. 930-935· 0 citations· 14 references
Abstract
Multimodal Sentiment Analysis (MSA) aims to recognize affective information by jointly exploiting signals from multiple modalities. Among textual, acoustic, and visual inputs, the textual modality usually conveys the primary semantic information associated with sentiment, whereas the other two modalities provide complementary nonverbal evidence. Based on this observation, this paper presents the Text-Guided Joint Interaction Network (TJINet), which promotes sufficient interaction between acoustic and visual information before introducing textual guidance. First, the features of each modality are independently encoded and transformed into compact representations. Next, the Gated Cross-Attention Joint Audio-Visual Interaction (JAVI-GCA) module performs bidirectional interaction between the acoustic and visual modalities and combines their complementary information into a joint audio-visual representation. Subsequently, the Language-guided Fusion Layer employs textual features as queries to selectively retrieve sentiment-related information from the previously fused audio-visual representation. The resulting multimodal representation is finally used to generate sentiment predictions. Experiments conducted on the CMU-MOSI and CH-SIMS datasets demonstrate that TJINet delivers better overall performance than several existing advanced methods.
A Text-Guided Hyper-modality Interaction Network (TGHIN) for multimodal sentiment analysis with differentiated feature encoding strategies for each modality and a Joint-Specific Fusion (JSF) module that enables the hyper-modality representation to refocus on the core information of each modality.
Kunxia Wang, RenLei Ding, YiHan Ge et al.· Signal, Image and Video Proc...· 0 citations
A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.
Chuhan Cheng, Hangcheng Wu, Junqiao Wang et al.· International Conference on...· 0 citations
Less obvious expressions like sarcasm and implicit sentiment are handled more effectively in this work, improving interpretation in multimodal sentiment analysis of social media data.
Prashant Adakane, Amit Gaikwad· international journal of eng...· 0 citations
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations
MagXCL, a unified framework designed to improve multimodal integration through more effective interaction between verbal and non-verbal modalities, is proposed, demonstrating the effectiveness of combining AMag with CrossCL to produce more accurate and robust multimodal sentiment predictions.
An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.
Bin Song, Wenjing Liu, Zhipeng Liang et al.· Signal, Image and Video Proc...· 0 citations