A prototype-as-prompt framework that maps audio–visual representations into a fixed set of multimodal sentiment prototypes that are used as soft prompts to guide the LLM in performing MSA and introduces a sentiment-aware prototype learning that explicitly binds multimodal prototypes with sentiment semantics.
MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.
Shanshan Lin, Yuesheng Wu, Chao Chen et al.· 0 citations
SentiLLM is proposed, a unified framework that leverages Semantic-Aligned Structural Abstraction to distill continuous raw signals into compact, semantically meaningful tokens and significantly improves discriminative performance with only a small number of trainable parameters.
Wei Chen, Junkai Li, Tongguan Wang et al.· 0 citations
This work proposes LLM-Augmented Prompt Learning for Multimodal Sentiment Analysis with Reward Adaptation (LAPM-RA), a unified framework integrating LLM-based sentiment-aware augmentation, reward-guided prompt selection, and context-aware multimodal fusion.
In multimodal sentiment analysis, textual, acoustic, and visual modalities often contain redundant and noisy information. Such information increases model complexity and weakens core sentiment representations, degrading accuracy and robustness. To address this issue, we propose CLIBN, a multimodal sentiment recognition network based on contrastive learning and information bottleneck. First, we design a sentimentintensity-aware contrastive learning strategy. It constructs positive and negative pairs according to sentiment intensity distances and assigns adaptive weights to different pairs, enabling the model to capture fine-grained sentiment differences. Second, we introduce a hierarchical information bottleneck module. It treats text as the primary modality and progressively integrates complementary cues from acoustic and visual modalities, while preserving task-relevant semantics and suppressing redundant information. Experimental results on CMU-MOSI and CMU-MOSEI show that CLIBN achieves superior performance. Specifically, Acc-2 reaches 87.8% and 86.7%, and F1-Score reaches 87.8% and 86.6% on the two datasets, respectively. These results demonstrate the effectiveness of CLIBN for multimodal sentiment representation learning.
Xu Meng, Yi Zhang, Yang Li· 2026 8th International Confe...· 0 citations
A multimodal solution that integrates audio and text information via cross-modal transformers, where text transcripts are automatically generated via an automatic speech recognition (ASR) tool and multiple text modalities are created via a cascaded architecture comprising cross-modal transformer blocks that integrate modalities one by one.
Andrei-George Durdun, V. Constantinescu, R. Ionescu· 0 citations
The framework first introduces learnable sentiment prototypes as semantic anchors to provide explicit sentiment-discriminative guidance for feature completion, and a gradient decoupling strategy is designed to separate the optimization paths of unimodal and multimodal objectives, preventing fusion gradients from interfering with unimodal encoders, thereby synergistically enhancing both discriminative representation learning and multimodal fusion.
Shan Tao, Haipeng Chen, Yu Liu et al.· Multimedia Systems· 0 citations