A Context-Aware Multimodal Transformer for the Affective Monitoring of Human Cognitive Behavior Using Facial and Speech Emotion Recognition
Abstract
Prolonged tracking of affective states empowers caregivers, educators, and medical staff to identify initial signs of distress. This is particularly vital for vulnerable individuals, such as the elderly and children, who may struggle to verbally express their emotional needs. This paper proposes a compact audio--visual system that links Facial Emotion Recognition (FER) with Speech Emotion Recognition (SER) and augments the fused representation with a context-aware multimodal Transformer. The visual branch is a Convolutional Neural Network (CNN) trained on FER2013, while the acoustic branch is a CNN--BiLSTM network driven by log-Mel spectrograms extracted from the RAVDESS corpus. Each branch outputs a distribution over emotion categories, and these distributions are combined through a single-parameter decision-level fusion rule based on a weighted sum. The fused scores are embedded and passed to a lightweight Transformer encoder that models temporal evolution and contextual dependencies across short interaction windows. Accuracy, macro precision, recall, and F1-score are the considered evaluation metrics, while class-wise analysis, ablation studies, and learning curves were also taken into account. On the datasets considered in this study, the joint audio–visual model with context-aware Transformer yields around a 5-7% absolute gain in accuracy over the best individual branch and a consistent improvement over a static fusion baseline. We further outline how such a network might serve as the core of monitoring tools for children and older adults, and note a number of practical, privacy, and ethical issues that would need to be addressed for deployment.