Skip to content

Exposing Weaknesses in Emotion Recognition in Conversations

Sep 2026 · 0 citations · 11 references
Computer Science

TL;DR

An LLM-as-Judge framework is introduced that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision, suggesting that standard single-label evaluation is therefore insufficient.

Abstract

Emotion Recognition in Conversations (ERC) aims to identify speakers'emotions in multi-turn dialogue. Accurate emotion recognition can support a wide range of applications, including empathetic conversational agents, mental health support, and educational technologies. While many recent approaches rely on task-specific fine-tuning, such models may exploit dataset-specific cues. A central yet rarely questioned assumption in ERC is that each utterance can be assigned a single unambiguous emotion label. To investigate this assumption, we study ERC using Large Language Models (LLMs) in a zero-shot setting while incorporating preceding conversational turns as context. We show that aggregate metrics mask systematic failures. Errors concentrate around utterances containing negations, exclamations, and interjections. This pattern is consistent across all evaluated models, suggesting limitations in the benchmarks rather than model-specific weaknesses. A controlled re-annotation study involving four human annotators supports this finding: strong agreement is observed in only 35 percent of cases, with neutral utterances dominating high-agreement instances, while many emotional categories fall into low-agreement regimes. These findings suggest that many apparent model errors reflect genuine annotation ambiguity rather than poor emotion understanding. Standard single-label evaluation is therefore insufficient. To address this limitation, we introduce an LLM-as-Judge framework that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision.

View source

Similar papers

Conference Open access Sep 2026

MindTracker: Unveiling Implicit Emotions in Long-Horizon Dialogues

This work constructs the MaskDialog dataset carefully curated from television drama and large language models, and proposes two LLM-based baseline approaches, i.e., One-shot Self-consistent Inference and Cascaded Multi-step Inference, and conducts comprehensive analyses on dialogue construction strategies and inference...

Zhi-Qiang Gao, Jing Han, Zhuo-Chu Wang et al. · 0 citations
Open access Sep 2026

Beyond Labels: Training Cognitive Empathy with Context-Rich Synthetic Dialogues

RelationalDialogues is introduced, a novel, fully synthetic dataset of 12,849 multi-turn dialogues designed to explicitly train perspective-taking, demonstrating that training on highly contextualized, metadata-driven synthetic data is an effective method for advancing LLMs from displaying superficial sympathy to engag...

Neema Owji, Cameron Buckner · 0 citations
#artificial intelligence Preprint Sep 2026

Finding Emotions Where They Belong: Rethinking Audio Emotion Recognition through Masked Temporal Affective Grounding

Audio emotion recognition (AER) typically assigns a single label to an entire recording, leaving the temporal scope of that label ambiguous when multiple speakers and affective events are present. We address this limitation by reformulating AER as a Temporal Affective Grounding (TAG) task that associates emotions with...

Abdelrahman Mohamed, Lars Kai Hansen, Zheng-Hua Tan · 0 citations

When Models Hear What They Expect: Diagnosing Prosodic Heuristics in Multimodal Sarcasm Detection

Sarcasm detection is addressed through sarcasm detection, evaluating Qwen2.5-Omni and Qwen3-Omni on Mandarin Chinese and English under five modality conditions that decompose the contributions of lexical content, vocal semantics, and prosodic structure to reveal a shared stereotype of expressive prosody.

Yong-Jian Chen, Pengfei Wei, Yiqun Sun et al. · 1 citation

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.