Skip to content

Aspect-guided dual-branch fusion network for multimodal aspect-based sentiment analysis

Aug 2026 · Signal, Image and Video Processing · Vol 20 · 0 citations · 43 references

TL;DR

An Aspect-guided dual-branch fusion network (ADFN) to enhance sentiment prediction by incorporating external knowledge and integrating coarse and fine information is proposed, which incorporates syntactic dependency information to complement and enrich the textual semantic representations.

View source

Similar papers

Open access Aug 2026

Image–Text Multimodal Sentiment Analysis with Large Model-Generated Descriptive Semantics and Difference-Aware Gated Fusion

Image–text multimodal sentiment analysis aims to integrate textual and visual information to comprehensively understand sentiment expressions in complex scenarios. However, existing methods focus on cross-modal feature interaction and fusion, and still have difficulty capturing effective sentiment cues in scenarios involving insufficient textual semantics, implicit visual affective cues, and inconsistent sentiment expressions between text and image. To address these issues, this paper proposes an image–text multimodal sentiment analysis method with large model-generated descriptive semantics and difference-aware gated fusion. Specifically, a large model generates semantic descriptions for image–text pairs, from which an enhanced semantic view is constructed to supplement implicit or insufficiently expressed sentiment cues in the original modalities. An original-enhanced dual-branch structure models the original image–text evidence and enhanced semantic evidence separately. To improve semantic consistency between the two branches, a cross-branch semantic alignment mechanism is introduced to reduce semantic shifts caused by enhanced information. In the fusion stage, difference-aware gated fusion and residual compensation are employed to adaptively balance branch contributions while preserving discriminative branch differences. Experimental results on the MVSA-Single and MVSA-Multiple datasets show that the proposed method improves performance in image–text multimodal sentiment classification, thereby validating the effectiveness of combining semantic enhancement with difference-aware modeling.

Hengyuan Zhang, Aizihaierjiang Yusufu, Jiang Liu et al. · 0 citations
Conference Jul 2026

Sentiment-Intensity-Aware Contrastive Learning with Hierarchical Information Bottleneck for Multimodal Sentiment Analysis

In multimodal sentiment analysis, textual, acoustic, and visual modalities often contain redundant and noisy information. Such information increases model complexity and weakens core sentiment representations, degrading accuracy and robustness. To address this issue, we propose CLIBN, a multimodal sentiment recognition network based on contrastive learning and information bottleneck. First, we design a sentimentintensity-aware contrastive learning strategy. It constructs positive and negative pairs according to sentiment intensity distances and assigns adaptive weights to different pairs, enabling the model to capture fine-grained sentiment differences. Second, we introduce a hierarchical information bottleneck module. It treats text as the primary modality and progressively integrates complementary cues from acoustic and visual modalities, while preserving task-relevant semantics and suppressing redundant information. Experimental results on CMU-MOSI and CMU-MOSEI show that CLIBN achieves superior performance. Specifically, Acc-2 reaches 87.8% and 86.7%, and F1-Score reaches 87.8% and 86.6% on the two datasets, respectively. These results demonstrate the effectiveness of CLIBN for multimodal sentiment representation learning.

Xu Meng, Yi Zhang, Yang Li · 0 citations
Open access Jul 2026

Label-guided data augmentation for cross-domain aspect-based sentiment analysis via enhanced affinity fusion

Aspect-Based Sentiment Analysis (ABSA) aims to identify opinion targets and determine the sentiment polarity expressed toward each aspect, which requires fine-grained modeling of aspect–opinion relations. Despite recent advances, cross-domain ABSA remains challenging due to structural mismatches across domains and the scarcity of high-quality labeled data in target domains. Existing methods often struggle to jointly address relational modeling errors and data sparsity, particularly under low-resource and cross-domain settings. To tackle these challenges, we propose a structure- and data-co-enhanced framework for cross-domain ABSA. At the model level, we introduce an Enhanced Affinity Fusion (EAF) module that explicitly strengthens aspect–opinion relational modeling by selectively integrating complementary attention mechanisms. Specifically, EAF combines biaffine attention to capture second-order interactions with syntax-aware attention to inject structural inductive bias, enabling robust modeling of long-distance dependencies without introducing excessive architectural complexity. At the data level, we propose Label-Guided Data Amplification (LGDA), which enhances supervision diversity and domain robustness through label-driven text expansion, hard sample mining, and domain-adaptive sampling. By jointly enhancing structural representation learning and training data supervision, the proposed framework effectively alleviates both aspect–opinion mismatches and cross-domain data sparsity. Extensive experiments on benchmark ABSA datasets demonstrate that our approach consistently outperforms strong baselines and achieves state-of-the-art performance in cross-domain scenarios. Ablation studies further validate the complementary contributions of EAF and LGDA.

Ningning Mao, Xuanliang Zhu, J. Wei et al. · 0 citations
Conference 2026

Hierarchical Global-Local Interaction and Refinement for Multimodal Sentiment Analysis

A Modality Dropout strategy is first introduced at the input stage to alleviate over-reliance on a sin-gle modality and improve robustness and the proposed Hierarchical Global-Local Interaction and Refinement framework for Multimodal Sentiment Analysis (HGLIR) is proposed.

yuanyuan zhou · 0 citations
#small language model Preprint Aug 2026

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.

Shanshan Lin, Yuesheng Wu, Chao Chen et al. · 0 citations