Multimodal Emotion Recognition System Using Facial Expressions and Speech Analysis
Abstract
Human communication involves emotion as an important issue. Because of emotion, people express and understand feelings, attitudes, and actions. However, effective automatic emotion recognition is still difficult to achieve when based on one single cue, as for instance facial expressions or voice may both present weak or ambiguous information. This study presents a Multimodal Emotion Recognition System that combines facial and speech information to improve emotion classification. The visual module uses deep learning to identify emotion-related facial patterns, while the speech module examines acoustic, spectral, and prosodic characteristics of voice signals. The prediction probabilities produced by the two modalities are integrated through a confidence-aware dynamic fusion mechanism, allowing the contribution of each source to change according to its estimated reliability. The proposed framework was evaluated on 2,600 held-out test samples using accuracy, precision, recall, F1-score, and specificity. The speech-only model achieved 76.50% accuracy, whereas the facial model achieved 82.40%. Combining both modalities through static fusion increased accuracy to 87.10%, while the proposed dynamic fusion method achieved the highest accuracy of 91.80%. The multimodal system also obtained 91.50% precision, 91.60% recall, 91.50% F1-score, and 98.27% specificity. Among the seven emotion categories, happy produced the highest F1-score at 95.90%, followed by neutral at 94.60%, surprise at 92.70%, angry at 90.90%, disgust at 89.70%, sad at 88.80%, and fear at 87.70%. In general, the results indicate that the merging of audio and visual information improves the identification of emotions and allows more accurate forecasting when there is a doubt of one means of communication.