Vision-Guided Text Prompt Tuning for Multimodal Sentiment Analysis
Vision-Guided Text Prompt Tuning (VG-TPT), which formulates visual-text sentiment modeling as controllable visual calibration of frozen text representations, consistently improves over text-only baselines and achieves competitive or superior performance compared with several full-modality methods.