Skip to content
Open access

Beyond Labels: Training Cognitive Empathy with Context-Rich Synthetic Dialogues

Sep 2026 · UF Journal of Undergraduate Research · Vol 28 · 0 citations

TL;DR

RelationalDialogues is introduced, a novel, fully synthetic dataset of 12,849 multi-turn dialogues designed to explicitly train perspective-taking, demonstrating that training on highly contextualized, metadata-driven synthetic data is an effective method for advancing LLMs from displaying superficial sympathy to engaging in genuine cognitive perspective-taking, a critical step for developing more emotionally intelligent conversational agents.

Abstract

As Large Language Models (LLMs) are increasingly deployed in emotionally sensitive applications like mental health support, the need for genuine artificial empathy has become critical. Existing conversational datasets often fail to instill deep cognitive empathy, instead promoting superficial affective mimicry. To address this gap, we introduce RelationalDialogues, a novel, fully synthetic dataset of 12,849 multi-turn dialogues designed to explicitly train perspective-taking. Using a 120-billion parameter LLM, we developed a data generation pipeline that grounds each conversation in structured metadata, including distinct speaker/listener backgrounds, situational stimuli, and relationship dynamics derived from a psychologically justified emotion taxonomy. A key feature of our methodology is blinding the "listener" agent to the explicit emotion label during generation, forcing it to infer emotional states from conversational context alone. We fine-tuned a Llama-3-8B-Instruct model on this corpus and evaluated it against its base counterpart using an LLM-as-a-Judge framework across three dimensions: Emotion Recognition, Perspective-Taking, and Emotional Contagion. Our fine-tuned model achieved a 51.05% win rate over the base model (45.53%) and demonstrated statistically significant improvements, particularly in Perspective-Taking (average score increase from 3.06 to 3.22) and Emotional Contagion (2.96 to 3.09). Granular analysis reveals the fine-tuned model excels at navigating nuanced, socially complex emotions like loneliness and jealousy, whereas the base model relies on pre-programmed templates for high-arousal emotions. These findings demonstrate that training on highly contextualized, metadata-driven synthetic data is an effective method for advancing LLMs from displaying superficial sympathy to engaging in genuine cognitive perspective-taking, a critical step for developing more emotionally intelligent conversational agents.

Read PDF

Similar papers

Conference Open access Sep 2026

MindTracker: Unveiling Implicit Emotions in Long-Horizon Dialogues

This work constructs the MaskDialog dataset carefully curated from television drama and large language models, and proposes two LLM-based baseline approaches, i.e., One-shot Self-consistent Inference and Cascaded Multi-step Inference, and conducts comprehensive analyses on dialogue construction strategies and inference...

Zhi-Qiang Gao, Jing Han, Zhuo-Chu Wang et al. · 0 citations
Preprint Oct 2026

EchoChat: Structured Cognitive Reasoning in Empathetic Spoken Dialogue

Empathetic spoken dialogue is a sophisticated cognitive process that requires not only recognizing emotions but also inferring a user's latent mental states to provide appropriate support. However, current SpeechLLMs often treat empathy as a direct input-to-response mapping, leading to"superficially warm"but emotionall...

Ding-Dong Wang, Shu-Jie Liu, Ya-Yue Deng et al. · 0 citations
Preprint Aug 2026

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models

EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations, and EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning and Group Relative Policy Optimization, are introduced, establishing a foundational framework for advan...

Junyu Wang, Si-Yuan Zhang, Peiyuan Jiang et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Exposing Weaknesses in Emotion Recognition in Conversations

An LLM-as-Judge framework is introduced that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision, suggesting that standard single-label evaluation is therefore insufficient.

Amir Ben Khalifa, Fanny Bezancon, B. Abdulrazak et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Text: LLM-Based Dimensional Emotion Evaluation in Multimodal Dialogue

Emotion recognition in conversation has been widely studied, but applying Large Language Models (LLMs) to continuous dimensional emotion evaluation in multimodal dialogue remains largely unexplored. We propose an LLM-based framework that performs discrete emotion recognition and Valence-Arousal-Dominance (VAD) dimensio...

Yu-Tong Hu, Jin-Ho D. Choi · 0 citations
Book Open access Oct 2026

Do Emotional Cues Matter? Exploring Support Strategy Selection in LLM-Based Supportive Conversations

Multimodal human–AI systems increasingly accept both text and speech, yet speech carries paralinguistic emotional cues that text does not. While prior work has evaluated the quality of large language model (LLM) responses, little is known about how vocal emotional cues reshape the support strategies an LLM selects and...

Wei-Yi Tian, Safak Dogan, Jie Meng · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.