Sep 2026· Companion Publication of the 28th International Conference on Multimodal Interaction· 0 citations· 27 references
Computer Science
TL;DR
This study forms continuous recognition of the Group Emotion at a one-second resolution, introduces the Mixed state, which captures the emotional divergence among participants in the group, and proposes a multimodal temporal framework that integrates audio and video information using a sliding-window context.
Abstract
To realize natural behavior in dialogue agents in multi-party dialogue scenarios, it is important to understand group emotion such as valence and arousal as a whole. Most prior work addressed this task at the utterance level or using a coarse-grained time window, which is not sufficient to capture emotional dynamics. In this study, we formulate continuous recognition of the Group Emotion at a one-second resolution. Moreover, we also introduce the Mixed state, which captures the emotional divergence among participants in the group. We constructed a dataset with frame-level soft labels based on the TEIDAN corpus and propose a multimodal temporal framework that integrates audio and video information using a sliding-window context. Experimental results demonstrate that the temporal Transformer outperforms simple baselines and shows stronger temporal agreement with the ground-truth labels than the LLM-based model. The effect of context length is limited, whereas audio-visual input outperforms either unimodal input on the continuous-label metrics. Additionally, our analysis shows larger Group Emotion recognition errors in intervals with high Mixed values, exposing emotional divergence as a key challenge for group emotion recognition.
ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions, is introduced.
Arthur Peuvot, Romaric Besançon, Gaël de Chalendar et al.· 0 citations
EmotionDialogCN is introduced, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication and achieves an emotion distribution deviation from real human emotion statistics and consistent subject framing, translating into stable unimodal and multimodal performance across acousti...
Transformer-GAT is proposed, a hybrid framework that combines Transformer and the Graph Attention Network to enable cross-modal emotion understanding and effectively integrates multimodal features, balances global and local contexts, and provides deeper emotional insights, offering new directions for multimodal emotion...
A unified post-training framework that equips pretrained text-to-speech models with natural-language control over segment-level emotion and duration is proposed, highlighting post-training as a practical approach to extending existing speech synthesis models.