ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions, is introduced.
Abstract
Detecting emotions is necessary for building systems that can accurately and adaptively interact with humans. Speech Emotion Recognition (SER) has become an important research focus to develop intelligent spoken interfaces. However, most studies predict emotions at the utterance level, ignoring the conversational context, along with the emotional flow and speaker interactions it carries. In this paper, we introduce ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions. To evaluate the robustness of this method, we conducted experiments on datasets spanning diverse emotionally expressive styles and contexts. ACERT outperforms current state-of-the-art (SOTA) approaches on IEMOCAP, establishes the first context-aware benchmark on SAFE, and obtains strong results on MELD for unweighted, class-balanced metrics. Ablation studies show that ACERT's gains come from emotional and conversational continuity, rather than from speaker identity or acoustic conditions.
An LLM-as-Judge framework is introduced that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision, suggesting that standard single-label evaluation is therefore insufficient.
Amir Ben Khalifa, Fanny Bezancon, B. Abdulrazak et al.· 0 citations
Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted...
DSSM-CRF, an audio-only architecture that explicitly separates cross-speaker contextual influence and within-speaker emotion evolution, is proposed and matched controls demonstrate complementary gains from speaker-wise factorization and CRF modeling.
This study forms continuous recognition of the Group Emotion at a one-second resolution, introduces the Mixed state, which captures the emotional divergence among participants in the group, and proposes a multimodal temporal framework that integrates audio and video information using a sliding-window context.
Soma Iwata, K. Inoue, Mu-Yun Wu et al.· Companion Publication of the...· 0 citations
Audio emotion recognition (AER) typically assigns a single label to an entire recording, leaving the temporal scope of that label ambiguous when multiple speakers and affective events are present. We address this limitation by reformulating AER as a Temporal Affective Grounding (TAG) task that associates emotions with...
Abdelrahman Mohamed, Lars Kai Hansen, Zheng-Hua Tan· 0 citations
Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensio...
Yu-Tong Hu, Jin-Ho D. Choi· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.