Speech Emotion Recognition and Algorithmic Bias
Abstract
Although speech emotion recognition (SER) systems have been integrated into many sectors, including human-computer interaction, health, and educational technologies, concerns about their robustness and fairness persist, particularly across datasets, features, model architecture, and training strategies. This paper explores the impact of different SER design choices, particularly biases, across datasets, features, model architecture, and training strategies. The study uses IEMOCAP and RAVDESS as benchmark corpora. It compares a support vector machine (SVM) with handcrafted acoustic features to a convolutional and recurrent neural network (CRNN) model that operates on a spectrogram, as well as a fine-tuned self-supervised speech encoder. To the best of our knowledge, this is the first study to jointly examine cross-corpus generalization, fairness-aware training, and subgroup disparities by gender and corpus within a unified SER framework. The work explores the effects of fairness-centric training, data augmentation, and class imbalance on performance and subgroup disparity (gender and corpus). The results indicate that,, for deep models, large cross-corpus performance drops (10–30 percentage points) and emotion-specific confusion spersisth is an improvement over the SVM baseline (~65% macro F1) in spein A class-balanced training mechanism and data augmentation are effective in enhancing recognition of minority emotions, raising F1 scores for emotions such as fear and disgust from below 0.40 to 0.52–0.58, while fairness-centric loss mechanisms are effective at reducing the performance gap (from 4–8% to 2–3%) at the cost of global accuracy (1–3% reduction).