Skip to content
Conference

Multimodal Emotion Recognition with Emotion-Specific Cross-Modal Attention Blocks

Aug 2026 · Moratuwa Engineering Research Conference · pp. 802-807 · 0 citations · 13 references

Abstract

Multimodal emotion recognition is increasingly important for healthcare, education, and human-computer interaction. However, many existing systems learn a single shared representation for all emotions, which can blur subtle class-specific cues. This paper proposes an emotion-specific multimodal architecture that combines Wav2Vec speech embeddings, prosody features, and BERT text embeddings through emotion-specific cross-modal attention blocks. Each emotion is assigned a dedicated attention block with its own query projection, while the acoustic modalities are merged using a learnable gated fusion mechanism. In this way, the model learns separate discriminative representations for angry, sad, happy, and neutral emotions rather than relying on a single shared attention path. Experiments on the IEMOCAP corpus show that the proposed model achieves 82.12% unweighted accuracy, outperforming baseline methods. The results indicate that emotion-aware cross-modal attention improves both robustness and class separability in multimodal emotion recognition.

View source

Similar papers

Aug 2026

Bidirectional joint cross-attention framework for transformer based audio–visual emotion recognition

Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.

Arman Sajjadi, M. Nekou, Sayna Sarvar et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition

Transformer-GAT is proposed, a hybrid framework that combines Transformer and the Graph Attention Network to enable cross-modal emotion understanding and effectively integrates multimodal features, balances global and local contexts, and provides deeper emotional insights, offering new directions for multimodal emotion...

Jia-Qi Qiao, Yifan Lyu, Xiu-Juan Xu · 0 citations
Preprint Sep 2026

Multimodal Emotion Recognition in Conversations via Class-Wise Adaptive Modality Fusion and Affective Geometry

Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational context and emotional dynamics. We extend the Self-Distillation Transformer architecture for ERC with appearance+geometry visual representations, class-wise adaptive modality...

Oriol Marín, Roger Marí, Gloria Haro et al. · 0 citations
Open access 2026

Multimodal Emotion Recognition in Urdu through Late Fusion of Fine-Tuned Speech and Text Representations

This study proposes a multimodal deep learning framework for Urdu emotion recognition by integrating speech and text modalities that surpasses the existing UMEDNet benchmark, demonstrating the effectiveness of transformer-based feature extraction and multimodal late fusion for Urdu emotion recognition.

Muhammad Sheraz, Adil Majeed, Shehzad Khalid et al. · 0 citations
Open access Aug 2026

Bimodal Speech Emotion Recognition Using a Hybrid CNN-LSTM Architecture with Sentiment Fusion

A bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture and the Multimodal EmotionLines Dataset is proposed, suggesting that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER.

Tze-Syn Yap, Lee-Yeng Ong · 0 citations
Sep 2026

Mamba-CrossMod: a multimodal affective analysis framework based on selective state space model

The proposed Mamba-CrossMod is a novel multimodal feature fusion framework that introduces Mamba-ATT, an enhanced attention mechanism based on a selective state-space model for capturing long-range dependencies with theoretically linear complexity.

Yi-Wen Tong, Jing Mu, Wen-Xin Chang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.