Intelligent Audio-based Emotion Recognition in Speech by Deep Learning and Feature Engineering Techniques
This work introduces ExpressNet, an optimum Multi-Layer Perceptron (MLP)-based SER model aimed to solve issues by leveraging a wide range of prosodic and spectral qualities incorporating Mel-Frequency Cepstral Coefficients (MFCCs), spectral contrast, and pitch variations.