Skip to content
Preprint

Vision-Guided Text Prompt Tuning for Multimodal Sentiment Analysis

Sep 2026 · 0 citations · 27 references
Computer Science

TL;DR

Vision-Guided Text Prompt Tuning (VG-TPT), which formulates visual-text sentiment modeling as controllable visual calibration of frozen text representations, consistently improves over text-only baselines and achieves competitive or superior performance compared with several full-modality methods.

Abstract

Multimodal sentiment analysis requires effective modeling of both verbal semantics and non-verbal affective cues. A central challenge is to calibrate text-centered sentiment understanding with visual facial evidence in a controlled, adaptive, and parameter-efficient manner. Text usually serves as the semantic anchor, whereas visual cues provide complementary evidence for ambiguous or implicit expressions; however, indiscriminate fusion may introduce visual noise and distort textual semantics. Moreover, fully fine-tuning large visual and textual encoders is costly and prone to overfitting on limited and scenario-dependent MSA benchmarks. To address these issues, we propose Vision-Guided Text Prompt Tuning (VG-TPT), which formulates visual-text sentiment modeling as controllable visual calibration of frozen text representations. VG-TPT injects visual affective cues into a frozen BERT encoder through layer-wise adaptive prompts, rather than relying on late-stage feature fusion or full backbone tuning. A co-guided router composes prompts from a trainable prompt bank according to both the evolving text state and the visual guidance feature, enabling sample-specific and layer-specific modulation. Experiments on CMU-MOSEI and CMU-MOSI show that VG-TPT consistently improves over text-only baselines and achieves competitive or superior performance compared with several full-modality methods, while updating only 2.4M trainable parameters. The code is available at https://github.com/ma-tubu/VG-TPT.

View source

Similar papers

#small language model Preprint Aug 2026

Multi-Granularity Sentiment Integration for LLM-Based Multimodal Sentiment Analysis

MGSI first encodes audio and visual streams at short-, medium-, and long-range temporal scales, preserving both local variations and global affective trends, and applies polarity- and intensity-aware enhancement to better handle ambiguous and near-neutral samples.

Shanshan Lin, Yuesheng Wu, Chao Chen et al. · 0 citations
Open access Sep 2026

Label-Aware Entropic Distributional Contrastive Alignment for Multimodal Sentiment Analysis

Multimodal sentiment analysis integrates text, audio, and visual signals to infer affective states. However, sentiment information shared across modalities is often entangled with modality-specific variation, and existing representation learning methods do not fully exploit the relations encoded by continuous sentiment...

Meng-Yao Wang, Xiu-Yang Meng, Chun-Ling Wang · 0 citations
Open access Aug 2026

Image–Text Multimodal Sentiment Analysis with Large Model-Generated Descriptive Semantics and Difference-Aware Gated Fusion

Experimental results on the MVSA-Single and MVSA-Multiple datasets show that the proposed method improves performance in image–text multimodal sentiment classification, thereby validating the effectiveness of combining semantic enhancement with difference-aware modeling.

Hengyuan Zhang, Aizihaierjiang Yusufu, Jiang Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MVFA: A Multi-View Text-Guided Multimodal Fusion LLM Adapter for Sentiment Analysis and Emotion Recognition

The multi-view text-guided multimodal fusion adapter (MVFA) is proposed, a parameter-efficient framework that augments frozen LLMs with strong multimodal reasoning capability and achieves state-of-the-art performance on key metrics while updating only a small fraction of parameters.

Peng-Fei Shao, Ji-Sheng Dang, Jia-Wen Fang et al. · 0 citations
Review Open access Aug 2026

Parameter-Efficient Adaptation and Benchmarking of Large Vision-Language Models for Multimodal Aspect-Based Sentiment Analysis

Multimodal aspect-based sentiment analysis (MABSA) predicts the sentiment expressed toward a target aspect by jointly using textual and visual information, supporting fine-grained opinion analysis in product reviews, brand monitoring, and customer feedback. However, existing approaches remain sensitive to irrelevant vi...

Ismail Ifakir, E. Nfaoui, Abderrahim Zannou · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.