Skip to content

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement

Jul 2026 · arXiv.org · Vol abs/2607.21061 · 1 citation · 88 references
Computer Science

TL;DR

This work introduces Emotion Statement Judgement (ESJ), a statement-verification formulation that preserves the expressiveness of the input space while constraining outputs to discriminative judgements, and builds EmObserver, an emotion-oriented MLLM optimized on ESJ through an elaborate multi-stage recipe.

Abstract

Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases. We attribute this gap to a structural mismatch between conventional AICA paradigms and the open-ended, instruction-driven nature of MLLMs, where further analysis reveals four major limitations: omission of plausible responses, limited emotion taxonomies, neglect of contextual factors, and labor-intensive annotation. To overcome these barriers, we introduce Emotion Statement Judgement (ESJ), a statement-verification formulation that preserves the expressiveness of the input space while constraining outputs to discriminative judgements. We further develop INSETS, a labor-efficient pipeline that instantiates ESJ at scale by constructing INSETS-462k and supporting MVEI, a rigorously refined benchmark spanning sentiment polarity, emotion interpretation, scene context, and perception subjectivity. Beyond evaluation, we build EmObserver, an emotion-oriented MLLM optimized on ESJ through an elaborate multi-stage recipe. Extensive evaluation of broad-spectrum MLLMs on MVEI reveals fine-grained insights into current artificial visual emotional intelligence, while experiments on multiple AICA benchmarks demonstrate the accuracy, generalization, and reasoning faithfulness of EmObserver. Collectively, these results establish ESJ as a practical formulation, MVEI as a comprehensive benchmark, and EmObserver as an advanced baseline for advancing MLLM-oriented visual emotional intelligence. Code will be released at: https://github.com/wdqqdw/EmObserver.

View source

Similar papers

Preprint Aug 2026

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models

EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations, and EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning and Group Relative Policy Optimization, are introduced, establishing a foundational framework for advan...

Junyu Wang, Si-Yuan Zhang, Peiyuan Jiang et al. · 1 citation
#human-computer interacti... Preprint Aug 2026

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

This paper introduces OneEmo, a unified affective generalist capable of mastering emotion perception, comprehension, and interaction, and proposes Emo-Chord, a novel reinforcement learning strategy that stabilizes optimization through unified multi-task reward allocation.

Jiahao Huang, Zheng Lian, Jingyi Zhang et al. · 0 citations
Preprint Aug 2026

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotional perception of a given image and explains the underlying rationale. Although related benchmarks ex...

Xiao-Lin Chen, Xue-Meng Song, Wenhao Shi et al. · 0 citations
Open access Sep 2026

CogSent: Cognition-Driven Multimodal Sentiment Analysis Through Fast–Slow Thinking

Sentiment analysis of multimodal social media data is of great importance, not only for recognizing objective information but also for capturing subjective emotional states. While single-modal sentiment analysis has achieved notable progress, existing multimodal approaches still face two key challenges: (1) inadequate...

Guo-Guo Ye, Qi-Qi Chen, Li-Qi Yan et al. · 0 citations
Open access Aug 2026

Large Language Model-Assisted Distillation–Fusion Framework for Visual Emotion Recognition

A large language model-assisted distillation–fusion framework (VERLADF) is proposed, which introduces emotion instruction data generated by GPT to fine-tune a VLM, thereby enhancing its emotional semantic understanding capability and adaptively fuses predictions from the instruction-tuned VLM and the distillation modul...

Yujun Ma, Yun-Jie Zeng, Zhi-Yuan Chen et al. · 0 citations
Preprint Aug 2026

NTDH: Complex Reasoning for Comprehensive Affective Analysis

Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is context-dependent, requiring conflicting cues to be reconciled rather than mapped directly to labels. Existing methods learn this mapping d...

Tianlei Zhu, Zhiwei Liu, Yu-Yan Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.