Skip to content

Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition

Sep 2026 · 0 citations · 28 references
Computer Science

TL;DR

Transformer-GAT is proposed, a hybrid framework that combines Transformer and the Graph Attention Network to enable cross-modal emotion understanding and effectively integrates multimodal features, balances global and local contexts, and provides deeper emotional insights, offering new directions for multimodal emotion computing.

Abstract

Multimodal emotion recognition is a key research area in affective computing, with applications in sentiment analysis, intelligent customer service, and human-computer interaction. However, existing methods often rely on single-modal features or simple multimodal fusion, failing to capture the synergy between global and local contexts, which limits model performance and emotion understanding. To address this challenge, we propose Transformer-GAT, a hybrid framework that combines Transformer and the Graph Attention Network to enable cross-modal emotion understanding. The Transformer is used to capture global semantic information, while the Graph Attention Network is employed to model fine-grained relationships between modalities, thereby enhancing the representation of emotional features. Experiments on the IEMOCAP and MELD datasets show that our model achieves weighted F1 scores of 72.45% and 77.37%, outperforming state-of-the-art methods. These results demonstrate that Transformer-GAT effectively integrates multimodal features, balances global and local contexts, and provides deeper emotional insights, offering new directions for multimodal emotion computing.

View source

Similar papers

Conference Aug 2026

Multimodal Emotion Recognition with Emotion-Specific Cross-Modal Attention Blocks

Multimodal emotion recognition is increasingly important for healthcare, education, and human-computer interaction. However, many existing systems learn a single shared representation for all emotions, which can blur subtle class-specific cues. This paper proposes an emotion-specific multimodal architecture that combin...

Gnanaseelan Dharshika, A. Ramanan · 0 citations
Sep 2026

Mamba-CrossMod: a multimodal affective analysis framework based on selective state space model

The proposed Mamba-CrossMod is a novel multimodal feature fusion framework that introduces Mamba-ATT, an enhanced attention mechanism based on a selective state-space model for capturing long-range dependencies with theoretically linear complexity.

Yi-Wen Tong, Jing Mu, Wen-Xin Chang et al. · 0 citations
Aug 2026

Bidirectional joint cross-attention framework for transformer based audio–visual emotion recognition

Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.

Arman Sajjadi, M. Nekou, Sayna Sarvar et al. · 0 citations
Open access 2026

Multimodal Emotion Recognition in Urdu through Late Fusion of Fine-Tuned Speech and Text Representations

This study proposes a multimodal deep learning framework for Urdu emotion recognition by integrating speech and text modalities that surpasses the existing UMEDNet benchmark, demonstrating the effectiveness of transformer-based feature extraction and multimodal late fusion for Urdu emotion recognition.

Muhammad Sheraz, Adil Majeed, Shehzad Khalid et al. · 0 citations
Open access Aug 2026

Graph-Based Adaptive Multimodal Transformer with Temporal Emotion Memory for Robust Conversational Emotion Recognition

Conversational emotion recognition has become an essential research area for developing intelligent human–computer interaction systems capable of understanding contextual emotional dynamics from multimodal information. Existing transformer-based approaches mainly focus on local contextual dependencies while often overl...

S. V · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.