Skip to content
Book Open access

Multimodal Temporal Modeling for Continuous Group Emotion Recognition in Multi-party Dialogues

Sep 2026 · Companion Publication of the 28th International Conference on Multimodal Interaction · 0 citations · 27 references
Computer Science

TL;DR

This study forms continuous recognition of the Group Emotion at a one-second resolution, introduces the Mixed state, which captures the emotional divergence among participants in the group, and proposes a multimodal temporal framework that integrates audio and video information using a sliding-window context.

Abstract

To realize natural behavior in dialogue agents in multi-party dialogue scenarios, it is important to understand group emotion such as valence and arousal as a whole. Most prior work addressed this task at the utterance level or using a coarse-grained time window, which is not sufficient to capture emotional dynamics. In this study, we formulate continuous recognition of the Group Emotion at a one-second resolution. Moreover, we also introduce the Mixed state, which captures the emotional divergence among participants in the group. We constructed a dataset with frame-level soft labels based on the TEIDAN corpus and propose a multimodal temporal framework that integrates audio and video information using a sliding-window context. Experimental results demonstrate that the temporal Transformer outperforms simple baselines and shows stronger temporal agreement with the ground-truth labels than the LLM-based model. The effect of context length is limited, whereas audio-visual input outperforms either unimodal input on the continuous-label metrics. Additionally, our analysis shows larger Group Emotion recognition errors in intervals with high Mixed values, exposing emotional divergence as a key challenge for group emotion recognition.

Read PDF

Similar papers

Preprint Aug 2026

EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue

EmotionDialogCN is introduced, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication and achieves an emotion distribution deviation from real human emotion statistics and consistent subject framing, translating into stable unimodal and multimodal performance across acousti...

Yi Zheng, Yifan Xu, Yan Zhou et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition

Transformer-GAT is proposed, a hybrid framework that combines Transformer and the Graph Attention Network to enable cross-modal emotion understanding and effectively integrates multimodal features, balances global and local contexts, and provides deeper emotional insights, offering new directions for multimodal emotion...

Jia-Qi Qiao, Yifan Lyu, Xiu-Juan Xu · 0 citations
Preprint Sep 2026

Post-Training Zero-Shot TTS for Fine-Grained Emotion and Duration Control via Natural Language

A unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration is proposed, highlighting post-training as a practical approach to extending existing speech synthesis models.

Lian-Ru Gao, Yu-Jie Guo, Yong Qin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.