Skip to content
Open access

Bimodal Speech Emotion Recognition Using a Hybrid CNN-LSTM Architecture with Sentiment Fusion

Aug 2026 · Future Internet · 0 citations · 21 references

TL;DR

A bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture and the Multimodal EmotionLines Dataset is proposed, suggesting that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER.

Abstract

Speech emotion recognition (SER) is a fundamental task in affective computing; however, traditional unimodal approaches often struggle to capture the complex emotional cues present in spontaneous conversational speech. Bimodal frameworks that integrate acoustic and textual information have therefore emerged to provide complementary semantic and acoustic representations. This study proposes a bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture. Using the Multimodal EmotionLines Dataset (MELD), the framework combines temporal acoustic features, statistical acoustic features, and predicted textual sentiment. Experimental results indicate that the proposed model achieves reliable recognition of majority emotion classes but exhibits limited performance on underrepresented minority classes due to severe class imbalance. To better understand the contribution of each modality, feature sufficiency and feature necessity analyses were conducted. Furthermore, an evaluation of alternative fusion strategies showed that the expressive attention networks did not provide meaningful performance improvements over simple feature concatenation. These findings suggest that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER, highlighting the importance of addressing data imbalance before pursuing more sophisticated multimodal architectures.

Read PDF

Similar papers

Conference Aug 2026

Speech Sentiment and Emotion Recognition: A Systematic Survey of Acoustic Features, Deep Models, and Transformer Architectures

Speech-based sentiment and emotion recognition is essential for affect-aware human-computer interaction, yet existing studies vary considerably in acoustic representations, learning architectures, label spaces, datasets, and evaluation protocols. This systematic survey reviews recent advances in speech emotion recognit...

M. Laxmi, N. Banu · 0 citations
Conference Aug 2026

Speech Emotion Recognition Using Transformer-Based Architectures with Self-Attention Mechanisms

Speech Emotion Recognition (SER) has become a key aspect in human-computer interaction, and affective computing, yet, current methods are faced with the challenge of modeling long-range context and fine-grain emotional expressions in speech signals. This paper has countered these shortcomings, giving a Transformer-base...

Dalphin Mary F, Binu Siva Singh S. K · 0 citations
Conference Aug 2026

Multimodal Emotion Recognition with Emotion-Specific Cross-Modal Attention Blocks

Multimodal emotion recognition is increasingly important for healthcare, education, and human-computer interaction. However, many existing systems learn a single shared representation for all emotions, which can blur subtle class-specific cues. This paper proposes an emotion-specific multimodal architecture that combin...

Gnanaseelan Dharshika, A. Ramanan · 0 citations
Open access Aug 2026

Learning Adaptive Cross-Modal Interactions for Multimodal Sentiment Analysis

A framework for learning adaptive cross-modal interactions for multimodal sentiment analysis that consistently outperforms previous methods and enhances multimodal representation capability for sentiment classification is proposed.

Chuhan Cheng, Hangcheng Wu, Jun-Qiao Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition

Transformer-GAT is proposed, a hybrid framework that combines Transformer and the Graph Attention Network to enable cross-modal emotion understanding and effectively integrates multimodal features, balances global and local contexts, and provides deeper emotional insights, offering new directions for multimodal emotion...

Jia-Qi Qiao, Yifan Lyu, Xiu-Juan Xu · 0 citations
Conference Open access Sep 2026

A Multimodal Speech Emotion Recognition Framework for Malayalam through Audio-Text Fusion

Speech emotion recognition (SER) is used in many domains, such as translation, intelligent assistants, healthcare monitoring, large language models, and human-computer interaction. Emotion recognition in Malayalam, however, remains challenging because of the language's rich morphological structure. This work introduces...

Athira Raj, Christy James Jose, K. Biju · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.