Skip to content
Open access

Intelligent Audio-based Emotion Recognition in Speech by Deep Learning and Feature Engineering Techniques

Aug 2026 · International journal of computer information systems and industrial management applications · 0 citations

TL;DR

This work introduces ExpressNet, an optimum Multi-Layer Perceptron (MLP)-based SER model aimed to solve issues by leveraging a wide range of prosodic and spectral qualities incorporating Mel-Frequency Cepstral Coefficients (MFCCs), spectral contrast, and pitch variations.

Abstract

Speech Emotion Recognition (SER) is now progressively vital for many practical uses including virtual assistants, customer service, and healthcare monitoring as well as for SER systems still suffer with environmental noise, speaker variability, and cross-lingual adaptation that affect their accuracy and generalizing power even if they have made tremendous progress. This work introduces ExpressNet, an optimum Multi-Layer Perceptron (MLP)-based SER model aimed to solve these issues by leveraging a wide range of prosodic and spectral qualities incorporating Mel-Frequency Cepstral Coefficients (MFCCs), spectral contrast, and pitch variations. ReLU activation and a softmax output layer allow the model to classify six emotional states: anger, disgust, fear, happiness, neutral, and sad using a deep learning architecture. We assess ExpressNet using the CREMA-D dataset and achieve a test accuracy of 92.97%, above the results of previous state-of-the-art approaches. Particularly real-time applications gain from the method since it helps to blend high classification accuracy with computing efficiency. Our work emphasizes the need of applying deep learning methods with enhanced feature engineering to improve SER performance. Furthermore, we show a thorough assessment over numerous benchmark datasets to show the power and applicability capacity of the model in many different settings. Apart from being better than other options, ExpressNet is a consistent choice for use in real-world settings since it has a low overfitting rate. This work advances emotional computing by providing a solid and scalable foundation for SER, which will enable further research in emotional recognition systems. The probable utilization of this technique extends to mental health monitoring and human-computer interaction because it demonstrates excellence at handling complex emotional patterns in voice signals. Extensive research into self-supervised learning and multimodal data integration and cross-lingual adaptation will improve the model's potential across multiple application scenarios.

Read PDF

Similar papers

Open access 2019

Speech Emotion Recognition using Convolutional Neural Networks and Recurrent Neural Networks with Attention Model

Speech emotion recognition is an upcoming subfield of automatic speech recognition that shares multiple similarities with mood recognition in music signals. Audio signals containing human speech are used as input to classification algorithms trained to recognize emotions in the form of audio features. This thesis outline...

G. Tomas, S. Weinzierl, Athanasios Lykartsis · 2 citations
Sep 2026

DeepAgeEmotionNet: A Multi-Task Deep Learning Framework for Speech-Based Age and Emotional State Assessment

The task of automatic speaker profiling based on speech signals becomes increasingly crucial in human-computer interaction, clinical voice assessment, and affective computing. Still, the tasks of speaker age group classification and emotion recognition are usually studied separately despite similar acoustic characteris...

R. Patole · 0 citations
Conference Open access Sep 2026

A Multimodal Speech Emotion Recognition Framework for Malayalam through Audio-Text Fusion

Speech emotion recognition (SER) is used in many domains, such as translation, intelligent assistants, healthcare monitoring, large language models, and human-computer interaction. Emotion recognition in Malayalam, however, remains challenging because of the language's rich morphological structure. This work introduces...

Athira Raj, Christy James Jose, K. Biju · 0 citations
Open access Sep 2026

Speech Emotion Recognition and Algorithmic Bias

Although speech emotion recognition (SER) systems have been integrated into many sectors, including human-computer interaction, health, and educational technologies, concerns about their robustness and fairness persist, particularly across datasets, features, model architecture, and training strategies. This paper expl...

Dheyaa Ahmed · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.