Skip to content
Conference

Text-Guided Joint Interaction Network for Multimodal Sentiment Analysis

Jun 2026 · 2026 IEEE 2nd International Conference on Electronics, Energy Systems and Power Engineering (EESPE) · pp. 930-935 · 0 citations · 14 references

Abstract

Multimodal Sentiment Analysis (MSA) aims to recognize affective information by jointly exploiting signals from multiple modalities. Among textual, acoustic, and visual inputs, the textual modality usually conveys the primary semantic information associated with sentiment, whereas the other two modalities provide complementary nonverbal evidence. Based on this observation, this paper presents the Text-Guided Joint Interaction Network (TJINet), which promotes sufficient interaction between acoustic and visual information before introducing textual guidance. First, the features of each modality are independently encoded and transformed into compact representations. Next, the Gated Cross-Attention Joint Audio-Visual Interaction (JAVI-GCA) module performs bidirectional interaction between the acoustic and visual modalities and combines their complementary information into a joint audio-visual representation. Subsequently, the Language-guided Fusion Layer employs textual features as queries to selectively retrieve sentiment-related information from the previously fused audio-visual representation. The resulting multimodal representation is finally used to generate sentiment predictions. Experiments conducted on the CMU-MOSI and CH-SIMS datasets demonstrate that TJINet delivers better overall performance than several existing advanced methods.

View source

Similar papers

Aug 2026

TGHIN: text-guided hyper-modality interaction network for multimodal sentiment analysis

A Text-Guided Hyper-modality Interaction Network (TGHIN) for multimodal sentiment analysis with differentiated feature encoding strategies for each modality and a Joint-Specific Fusion (JSF) module that enables the hyper-modality representation to refocus on the core information of each modality.

Kunxia Wang, RenLei Ding, YiHan Ge et al. · 0 citations
Open access Aug 2026

Learning Adaptive Cross-Modal Interactions for Multimodal Sentiment Analysis

A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.

Chuhan Cheng, Hangcheng Wu, Junqiao Wang et al. · 0 citations
#small language model Preprint Aug 2026

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.

Shanshan Lin, Yuesheng Wu, Chao Chen et al. · 0 citations
Aug 2026

MagXCL: enhanced multimodal adaptation gate and cross-modal contrastive learning for multimodal sentiment analysis

MagXCL, a unified framework designed to improve multimodal integration through more effective interaction between verbal and non-verbal modalities, is proposed, demonstrating the effectiveness of combining AMag with CrossCL to produce more accurate and robust multimodal sentiment predictions.

Duc-Duy Duong, Cam-Van Thi Nguyen, Duc-Trong Le · 0 citations
Aug 2026

Aspect-guided dual-branch fusion network for multimodal aspect-based sentiment analysis

An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.

Bin Song, Wenjing Liu, Zhipeng Liang et al. · 0 citations