MVFA: A Multi-View Text-Guided Multimodal Fusion LLM Adapter for Sentiment Analysis and Emotion Recognition
The multi-view text-guided multimodal fusion adapter (MVFA) is proposed, a parameter-efficient framework that augments frozen LLMs with strong multimodal reasoning capability and achieves state-of-the-art performance on key metrics while updating only a small fraction of parameters.