Aug 2026· Moratuwa Engineering Research Conference· pp. 802-807· 0 citations· 13 references
Abstract
Multimodal emotion recognition is increasingly important for healthcare, education, and human-computer interaction. However, many existing systems learn a single shared representation for all emotions, which can blur subtle class-specific cues. This paper proposes an emotion-specific multimodal architecture that combines Wav2Vec speech embeddings, prosody features, and BERT text embeddings through emotion-specific cross-modal attention blocks. Each emotion is assigned a dedicated attention block with its own query projection, while the acoustic modalities are merged using a learnable gated fusion mechanism. In this way, the model learns separate discriminative representations for angry, sad, happy, and neutral emotions rather than relying on a single shared attention path. Experiments on the IEMOCAP corpus show that the proposed model achieves 82.12% unweighted accuracy, outperforming baseline methods. The results indicate that emotion-aware cross-modal attention improves both robustness and class separability in multimodal emotion recognition.
Experiments show that the proposed framework outperforms unimodal baselines and existing fusion methods, indicating that the approach learns context-aware emotion representations well suited for accuracy-oriented audio–visual emotion recognition applications.
Arman Sajjadi, M. Nekou, Sayna Sarvar et al.· Signal, Image and Video Proc...· 0 citations
Transformer-GAT is proposed, a hybrid framework that combines Transformer and the Graph Attention Network to enable cross-modal emotion understanding and effectively integrates multimodal features, balances global and local contexts, and provides deeper emotional insights, offering new directions for multimodal emotion...
Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational context and emotional dynamics. We extend the Self-Distillation Transformer architecture for ERC with appearance+geometry visual representations, class-wise adaptive modality...
Oriol Marín, Roger Marí, Gloria Haro et al.· 0 citations
This study proposes a multimodal deep learning framework for Urdu emotion recognition by integrating speech and text modalities that surpasses the existing UMEDNet benchmark, demonstrating the effectiveness of transformer-based feature extraction and multimodal late fusion for Urdu emotion recognition.
Muhammad Sheraz, Adil Majeed, Shehzad Khalid et al.· Computer Modeling in Enginee...· 0 citations
A bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture and the Multimodal EmotionLines Dataset is proposed, suggesting that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER.
The proposed Mamba-CrossMod is a novel multimodal feature fusion framework that introduces Mamba-ATT, an enhanced attention mechanism based on a selective state-space model for capturing long-range dependencies with theoretically linear complexity.
Yi-Wen Tong, Jing Mu, Wen-Xin Chang et al.· Pattern Analysis and Applica...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.