A modified mobilenetV2 using selective kernels and coordinate attention for facial expression recognition
Abstract
Facial Expression Recognition (FER) plays an important role in human-computer interaction (HCI). However, most highly accurate FER models rely on complex and computationally heavy architectures that makes them unsuitable for low resource devices. This study proposes a modified MobileNetV2-based lightweight network that integrates Selective Kernel (SK) and Coordinate Attention (CA) modules to enhance feature extraction and spatial awareness while maintaining a low parameter count. The SK layer dynamically adapts receptive field sizes to capture both narrow and broad facial details. The CA module encodes both spatial and channel dependencies to help the network focus on emotion-specific regions such as the eyes, mouth, and eyebrows. The proposed model was trained and evaluated on two benchmark datasets, RAF-DB and FER-2013. Experimental results demonstrate that the model achieved 87.38% accuracy on RAF-DB and 70.39% accuracy on FER-2013 with only 1.07 million parameters and 0.687 GFLOPs. A Unity WebGL implementation also demonstrate real-time validation of the model’s ability.