2026· Computers, Materials & Continua· Vol 88, pp. 1-10· 0 citations· 36 references
TL;DR
TC-DSC is proposed, a text-centric hierarchical dual-stream interaction framework for incomplete multimodal sentiment analysis that achieves competitive performance and consistent improvements under both complete and incomplete settings.
Abstract
: Incomplete multimodal sentiment analysis has attracted increasing research interest in recent years. Existing methods attempt to recover missing modalities through generative reconstruction and text-enhanced fusion, but these approaches may be limited in preserving sentiment-relevant information and fully leveraging complementary and hierarchical cross-modal interactions, particularly under noisy or incomplete conditions. To address these challenges, we propose TC-DSC, a text-centric hierarchical dual-stream interaction framework for incomplete multimodal sentiment analysis. Rather than reconstructing raw signals, TC-DSC performs semantic alignment and consistency modeling in the feature space through structured interactions between a text-centric stream and auxiliary audio-visual streams. A multi-scale enhanced encoder is designed to improve the robustness of non-text modalities under noisy conditions. Furthermore, a hierarchical proxy layer enables bidirectional interaction, with the text modality serving as a semantic anchor to guide cross-modal alignment. A semantic distillation strategy is also incorporated to facilitate knowledge transfer in the feature space under modality missing. Extensive experiments on MOSI, MOSEI, and SIMS demonstrate that TC-DSC achieves competitive performance and consistent improvements under both complete and incomplete settings.
An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.
Bin Song, Wenjing Liu, Zhipeng Liang et al.· Signal, Image and Video Proc...· 0 citations
HAFT routes audio–visual interaction through bottleneck tokens with adaptive depth-wise gating before integrating text, jointly reducing attention cost and counteracting modality bias; the Cascaded Audio Feature Enhancement (CAFE) framework strengthens prosodic representations via multi-scale time–frequency extraction; and grouped projections, decoupled positional attention, and layer-wise parameter sharing compress the remaining overhead.
Qing Dong, Ting Lu, Xiujin Shi et al.· International journal of sof...· 0 citations
A Text-Guided Hyper-modality Interaction Network (TGHIN) for multimodal sentiment analysis with differentiated feature encoding strategies for each modality and a Joint-Specific Fusion (JSF) module that enables the hyper-modality representation to refocus on the core information of each modality.
Kunxia Wang, RenLei Ding, YiHan Ge et al.· Signal, Image and Video Proc...· 0 citations
A Modality Dropout strategy is first introduced at the input stage to alleviate over-reliance on a sin-gle modality and improve robustness and the proposed Hierarchical Global-Local Interaction and Refinement framework for Multimodal Sentiment Analysis (HGLIR) is proposed.
yuanyuan zhou· Poster Volume 0007 The 2026...· 0 citations
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations