Aug 2026· Signal, Image and Video Processing· Vol 20· 0 citations· 53 references
TL;DR
A Text-Guided Hyper-modality Interaction Network (TGHIN) for multimodal sentiment analysis with differentiated feature encoding strategies for each modality and a Joint-Specific Fusion (JSF) module that enables the hyper-modality representation to refocus on the core information of each modality.
MagXCL, a unified framework designed to improve multimodal integration through more effective interaction between verbal and non-verbal modalities, is proposed, demonstrating the effectiveness of combining AMag with CrossCL to produce more accurate and robust multimodal sentiment predictions.
A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.
Chuhan Cheng, Hangcheng Wu, Junqiao Wang et al.· International Conference on...· 0 citations
Less obvious expressions like sarcasm and implicit sentiment are handled more effectively in this work, improving interpretation in multimodal sentiment analysis of social media data.
Prashant Adakane, Amit Gaikwad· international journal of eng...· 0 citations
A Modality Dropout strategy is first introduced at the input stage to alleviate over-reliance on a sin-gle modality and improve robustness and the proposed Hierarchical Global-Local Interaction and Refinement framework for Multimodal Sentiment Analysis (HGLIR) is proposed.
yuanyuan zhou· Poster Volume 0007 The 2026...· 0 citations
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations
Multimodal Sentiment Analysis (MSA) integrates information from multiple modalities to infer sentiment. It faces two challenges: accurately modeling unimodal modalities and effectively fusing them. Existing graph neural network (GNN) based methods are limited by the over-smoothing problem caused by deep architectures, and thus typically adopt shallow structures for modality modeling. While these shallow structures can capture local information within a modality well, they struggle to capture global context. Meanwhile, current text-centric fusion methods do not explicitly align textual and non-textual modalities before fusion, which degrades fusion performance. To address these issues, we propose a modality-specific Graph Transformer with Prompt-aware Fusion (GTPF) framework. GTPF employs a Modality-Specific Graph Transformer (MSGT) architecture to explore both local and global information within modalities, enabling more accurate unimodal modeling. It also uses a Prompt-Aware Multimodal Fusion (PAMF) module that adopts prompt shifting to align textual and non-textual modalities before fusion, thereby enhancing text-centric fusion. We conduct extensive experiments on three MSA benchmarks. GTPF outperforms state-of-the-art methods across all metrics. On the CMU-MOSI dataset, GTPF achieves a relative improvement of up to 13.76% in mean absolute error (MAE) over the second-best model. On the CH-SIMS dataset, it achieves relative improvements of 6.01% in Pearson correlation (Corr) and 5.03% in five-class classification accuracy over the second-best model. These results validate the effectiveness of our Graph-Transformer co-design and prompt-aware fusion strategy.
Haolong Zheng, Yan Leng, Jia-Ning Wu et al.· Neural Networks· 0 citations