EXPLAINABLE REAL-TIME FACIAL EMOTION RECOGNITION: RESNET-50 VERSUS VISION TRANSFORMER WITH RUSSELL’S CIRCUMPLEX MODEL
Abstract
Facial emotion recognition systems must balance classification accuracy, computational efficiency, and model transparency in real-time settings. This study compares ResNet-50 and Vision Transformer (ViT-Base) on the FER2013 dataset and integrates both models into an explainable real-time prototype. Grayscale facial images were converted into three-channel 224 × 224 inputs. Transfer learning and weighted cross-entropy were applied to address class imbalance, and performance was evaluated through accuracy, precision, recall, F1-score, inference latency, and frame rate. MediaPipe extracted facial regions, while Grad-CAM and Russell’s valence-arousal circumplex supported visual and psychological interpretation. ViT-Base achieved 81.89% accuracy and 81.9% F1-score, exceeding ResNet-50 at 78.30% accuracy and 78.2% F1-score. This 3.59 percentage-point gain required greater computation: ViT used about 85 million parameters, 50–70 ms per frame, and 15–20 FPS, whereas ResNet-50 used about 25 million parameters, 30–40 ms per frame, and 25–30 FPS. Both models performed strongly on distinctive emotions, but sadness-neutral confusion remained dominant. ViT is preferable when accuracy is prioritized, while ResNet-50 is more practical for latency- and resource-constrained deployment.