Aug 2026· International journal of computer information systems and industrial management applications· 0 citations
TL;DR
This work introduces ExpressNet, an optimum Multi-Layer Perceptron (MLP)-based SER model aimed to solve issues by leveraging a wide range of prosodic and spectral qualities incorporating Mel-Frequency Cepstral Coefficients (MFCCs), spectral contrast, and pitch variations.
Abstract
Speech Emotion Recognition (SER) is now progressively vital for many practical uses including virtual assistants, customer service, and healthcare monitoring as well as for SER systems still suffer with environmental noise, speaker variability, and cross-lingual adaptation that affect their accuracy and generalizing power even if they have made tremendous progress. This work introduces ExpressNet, an optimum Multi-Layer Perceptron (MLP)-based SER model aimed to solve these issues by leveraging a wide range of prosodic and spectral qualities incorporating Mel-Frequency Cepstral Coefficients (MFCCs), spectral contrast, and pitch variations. ReLU activation and a softmax output layer allow the model to classify six emotional states: anger, disgust, fear, happiness, neutral, and sad using a deep learning architecture. We assess ExpressNet using the CREMA-D dataset and achieve a test accuracy of 92.97%, above the results of previous state-of-the-art approaches. Particularly real-time applications gain from the method since it helps to blend high classification accuracy with computing efficiency. Our work emphasizes the need of applying deep learning methods with enhanced feature engineering to improve SER performance. Furthermore, we show a thorough assessment over numerous benchmark datasets to show the power and applicability capacity of the model in many different settings. Apart from being better than other options, ExpressNet is a consistent choice for use in real-world settings since it has a low overfitting rate. This work advances emotional computing by providing a solid and scalable foundation for SER, which will enable further research in emotional recognition systems. The probable utilization of this technique extends to mental health monitoring and human-computer interaction because it demonstrates excellence at handling complex emotional patterns in voice signals. Extensive research into self-supervised learning and multimodal data integration and cross-lingual adaptation will improve the model's potential across multiple application scenarios.
Speech emotion recognition is an upcoming subfield of automatic speech recognition that shares multiple similarities with mood recognition in music signals. Audio signals containing human speech are used as input to classification algorithms trained to recognize emotions in the form of audio features. This thesis outline...
G. Tomas, S. Weinzierl, Athanasios Lykartsis· Proceedings of 2019 the 9th...· 2 citations
The obtained experimental results prove the superiority of the proposed hybrid representation over the single Wavelet and MFCC features, achieving the overall recognition accuracy of 99% and average accuracy of 94%.
M. Mohanty, R. Ram, Kumuda Sharma et al.· International Journal of Spe...· 0 citations
The task of automatic speaker profiling based on speech signals becomes increasingly crucial in human-computer interaction, clinical voice assessment, and affective computing. Still, the tasks of speaker age group classification and emotion recognition are usually studied separately despite similar acoustic characteris...
R. Patole· Natural Resources for Human...· 0 citations
Speech emotion recognition (SER) is used in many domains, such as translation, intelligent assistants, healthcare monitoring, large language models, and human-computer interaction. Emotion recognition in Malayalam, however, remains challenging because of the language's rich morphological structure. This work introduces...
Athira Raj, Christy James Jose, K. Biju· Conference Proceedings in Sc...· 0 citations
Although speech emotion recognition (SER) systems have been integrated into many sectors, including human-computer interaction, health, and educational technologies, concerns about their robustness and fairness persist, particularly across datasets, features, model architecture, and training strategies. This paper expl...