Skip to content

Enriching Speech Emotion Representations with Conversational Context

Sep 2026 · 0 citations · 30 references
Computer Science Engineering

TL;DR

ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions, is introduced.

Abstract

Detecting emotions is necessary for building systems that can accurately and adaptively interact with humans. Speech Emotion Recognition (SER) has become an important research focus to develop intelligent spoken interfaces. However, most studies predict emotions at the utterance level, ignoring the conversational context, along with the emotional flow and speaker interactions it carries. In this paper, we introduce ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions. To evaluate the robustness of this method, we conducted experiments on datasets spanning diverse emotionally expressive styles and contexts. ACERT outperforms current state-of-the-art (SOTA) approaches on IEMOCAP, establishes the first context-aware benchmark on SAFE, and obtains strong results on MELD for unweighted, class-balanced metrics. Ablation studies show that ACERT's gains come from emotional and conversational continuity, rather than from speaker identity or acoustic conditions.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Exposing Weaknesses in Emotion Recognition in Conversations

An LLM-as-Judge framework is introduced that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision, suggesting that standard single-label evaluation is therefore insufficient.

Amir Ben Khalifa, Fanny Bezancon, B. Abdulrazak et al. · 0 citations
#artificial intelligence Preprint Sep 2026

From Speech to Editable Concepts: Probing Emotion Recognition with Concept Bottleneck Models

Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent approaches combine multiple modalities, most commonly speech and text. Still, performance remains poor on many datasets. Large language models (LLMs) have therefore attracted...

He-Zhao Zhang, Thomas Hain · 0 citations
#machine learning Preprint Aug 2026

Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation

DSSM-CRF, an audio-only architecture that explicitly separates cross-speaker contextual influence and within-speaker emotion evolution, is proposed and matched controls demonstrate complementary gains from speaker-wise factorization and CRF modeling.

Guang-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen · 0 citations
Book Open access Sep 2026

Multimodal Temporal Modeling for Continuous Group Emotion Recognition in Multi-party Dialogues

This study forms continuous recognition of the Group Emotion at a one-second resolution, introduces the Mixed state, which captures the emotional divergence among participants in the group, and proposes a multimodal temporal framework that integrates audio and video information using a sliding-window context.

Soma Iwata, K. Inoue, Mu-Yun Wu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Finding Emotions Where They Belong: Rethinking Audio Emotion Recognition through Masked Temporal Affective Grounding

Audio emotion recognition (AER) typically assigns a single label to an entire recording, leaving the temporal scope of that label ambiguous when multiple speakers and affective events are present. We address this limitation by reformulating AER as a Temporal Affective Grounding (TAG) task that associates emotions with...

Abdelrahman Mohamed, Lars Kai Hansen, Zheng-Hua Tan · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue

Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensio...

Yu-Tong Hu, Jin-Ho D. Choi · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.