An LLM-as-Judge framework is introduced that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision, suggesting that standard single-label evaluation is therefore insufficient.
Abstract
Emotion Recognition in Conversations (ERC) aims to identify speakers'emotions in multi-turn dialogue. Accurate emotion recognition can support a wide range of applications, including empathetic conversational agents, mental health support, and educational technologies. While many recent approaches rely on task-specific fine-tuning, such models may exploit dataset-specific cues. A central yet rarely questioned assumption in ERC is that each utterance can be assigned a single unambiguous emotion label. To investigate this assumption, we study ERC using Large Language Models (LLMs) in a zero-shot setting while incorporating preceding conversational turns as context. We show that aggregate metrics mask systematic failures. Errors concentrate around utterances containing negations, exclamations, and interjections. This pattern is consistent across all evaluated models, suggesting limitations in the benchmarks rather than model-specific weaknesses. A controlled re-annotation study involving four human annotators supports this finding: strong agreement is observed in only 35 percent of cases, with neutral utterances dominating high-agreement instances, while many emotional categories fall into low-agreement regimes. These findings suggest that many apparent model errors reflect genuine annotation ambiguity rather than poor emotion understanding. Standard single-label evaluation is therefore insufficient. To address this limitation, we introduce an LLM-as-Judge framework that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision.
ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions, is introduced.
Arthur Peuvot, Romaric Besançon, Gaël de Chalendar et al.· 0 citations
EmoLASP's LLM pipeline demonstrates the potential advantages of using a reasoning approach to ensure emotion prediction consistency and to reduce both the cost of fine-tuning and the cost of prompting with long dialogue histories.
This work constructs the MaskDialog dataset carefully curated from television drama and large language models, and proposes two LLM-based baseline approaches, i.e., One-shot Self-consistent Inference and Cascaded Multi-step Inference, and conducts comprehensive analyses on dialogue construction strategies and inference...
Zhi-Qiang Gao, Jing Han, Zhuo-Chu Wang et al.· Proceedings of the Thirty-Fi...· 0 citations
RelationalDialogues is introduced, a novel, fully synthetic dataset of 12,849 multi-turn dialogues designed to explicitly train perspective-taking, demonstrating that training on highly contextualized, metadata-driven synthetic data is an effective method for advancing LLMs from displaying superficial sympathy to engag...
Neema Owji, Cameron Buckner· UF Journal of Undergraduate...· 0 citations
Audio emotion recognition (AER) typically assigns a single label to an entire recording, leaving the temporal scope of that label ambiguous when multiple speakers and affective events are present. We address this limitation by reformulating AER as a Temporal Affective Grounding (TAG) task that associates emotions with...
Abdelrahman Mohamed, Lars Kai Hansen, Zheng-Hua Tan· 0 citations
Sarcasm detection is addressed through sarcasm detection, evaluating Qwen2.5-Omni and Qwen3-Omni on Mandarin Chinese and English under five modality conditions that decompose the contributions of lexical content, vocal semantics, and prosodic structure to reveal a shared stereotype of expressive prosody.
Yong-Jian Chen, Pengfei Wei, Yiqun Sun et al.· 1 citation
With $2.1 million funding from Google.org, the open-source Public Transit Intelligence Hub will unify public transit monitoring, operations, and passenger communication.
Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.