Skip to content
Conference Open access

MindTracker: Unveiling Implicit Emotions in Long-Horizon Dialogues

Sep 2026 · Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence · 0 citations · 37 references

TL;DR

This work constructs the MaskDialog dataset carefully curated from television drama and large language models, and proposes two LLM-based baseline approaches, i.e., One-shot Self-consistent Inference and Cascaded Multi-step Inference, and conducts comprehensive analyses on dialogue construction strategies and inference behaviors.

Abstract

Affective computing has achieved notable success in recognizing explicit emotions from short, isolated dialogue segments. However, human emotions are often implicitly expressed, internally regulated, and dynamically evolve over extended interactions. Existing models struggle to disentangle internal emotional states from external expressions, and fail to capture the emotional inconsistency that emerges across long-horizon dialogues. To address this limitation, we introduce Emotional Inconsistency Analysis (EIA), a novel task that aims to identify and reason about discrepancies between implicit and explicit emotions over long-term conversational contexts. To support this task, we construct the MaskDialog dataset carefully curated from television drama and large language models (LLMs). We further propose two LLM-based baseline approaches, i.e., One-shot Self-consistent Inference and Cascaded Multi-step Inference, and conduct comprehensive analyses on dialogue construction strategies and inference behaviors. Extensive experiments across multiple mainstream LLMs reveal that EIA remains highly challenging, particularly in modeling implicit emotional trajectories and cross-turn inconsistency. Overall, EIA reframes emotion understanding from short-term recognition to longitudinal, implicit emotion tracking, with implications for dialogue systems and human–computer interaction.

Read PDF

Similar papers

Open access Sep 2026

Beyond Labels: Training Cognitive Empathy with Context-Rich Synthetic Dialogues

RelationalDialogues is introduced, a novel, fully synthetic dataset of 12,849 multi-turn dialogues designed to explicitly train perspective-taking, demonstrating that training on highly contextualized, metadata-driven synthetic data is an effective method for advancing LLMs from displaying superficial sympathy to engag...

Neema Owji, Cameron Buckner · 0 citations
#artificial intelligence Preprint Sep 2026

LongEmo: Towards Emotion Understanding and Reasoning in Long Videos

While recent Multimodal Large Language Models (MLLMs) have shown promise in affective computing, their reasoning capabilities are largely confined to short video clips with limited interactions. However, real-world emotions are not merely isolated instantaneous reactions but dynamic and cumulative processes deeply shap...

Shuo Zhang, Yifan Zhou, Han-Yu Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Exposing Weaknesses in Emotion Recognition in Conversations

An LLM-as-Judge framework is introduced that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision, suggesting that standard single-label evaluation is therefore insufficient.

Amir Ben Khalifa, Fanny Bezancon, B. Abdulrazak et al. · 0 citations
Book Open access Oct 2026

Unmasking Suppressed Emotion: A Multi-Agent Approach to Affective Dissonance

Detecting true felt emotions when speakers suppress or mask their internal state poses a fundamental challenge for affective computing systems. We study a specific form of affective dissonance in dyadic speech: utterances where a speaker’s self-reported emotion diverges from all external observer ratings, indicating th...

Cheng-Shiuan Lin, T. Eze, Da-Wei Xie et al. · 0 citations
Preprint Aug 2026

EmotionDialogCN: A Spontaneous Multimodal Dataset for Mandarin Emotional Dialogue

EmotionDialogCN is introduced, a large-scale audiovisual-emotional dataset designed to capture authentic face-to-face communication and achieves an emotion distribution deviation from real human emotion statistics and consistent subject framing, translating into stable unimodal and multimodal performance across acousti...

Yi Zheng, Yifan Xu, Yan Zhou et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.