Speech Sentiment and Emotion Recognition: A Systematic Survey of Acoustic Features, Deep Models, and Transformer Architectures
Abstract
Speech-based sentiment and emotion recognition is essential for affect-aware human-computer interaction, yet existing studies vary considerably in acoustic representations, learning architectures, label spaces, datasets, and evaluation protocols. This systematic survey reviews recent advances in speech emotion recognition and direct positive-negative-neutral sentiment classification, spectrogram and waveform representations, deep neural networks, self-supervised speech encoders, transformers, acoustic-linguistic fusion, and domain-adaptation methods. Relevant peer-reviewed studies were identified through major scholarly databases and comparatively analyzed according to input modality, feature representation, model family, prediction target, dataset, and within- and cross-corpus evaluation. The synthesis shows that multi-feature fusion, pretrained speech encoders, contextual language models, and adaptive learning strategies improve affect representation and contextual understanding. Direct speech-only sentiment classification, cross-corpus and cross-lingual generalization, computational efficiency, interpretability, and standardized evaluation remain insufficiently addressed. Based on these findings, the survey presents an objective-aligned taxonomy and a conceptual framework integrating complementary acoustic, semantic, contextual, and adaptive components. The study provides practical guidance for developing accurate, efficient, explainable, and generalizable speech-affect recognition systems.