Performance Evaluation of Hybrid CNN-Transformer Architectures for EEG-Based Data
Abstract
Electroencephalography (EEG) is widely used in analysing the emotions of humans. EEGs are also widely used in brain-computer interface (BCI), clinical diagnostic, and cognitive-monitoring applications etc. There are various deep learning approaches for decoding EEG signals.CNNs focus on learning local spatial and frequency-domain characteristics of EEG signals, whereas Transformer architectures capture broader temporal dependencies using self-attention mechanisms. This paper reviews recent literature that evaluates and compares the performance of CNN-based, Transformer-based, and hybrid CNN-Transformer models across EEG applications including motor imagery classification, emotion recognition, sleep-stage scoring, seizure detection, clinical abnormality screening, and inner-speech decoding. This study evaluated hybrid CNN-Transformer architectures—specifically the cnn_temporal_transformer and dual_spatial_temporal_transformer models—for EEG-based emotion recognition across the DEAP and SEED datasets, spanning arousal, valence, and multi-class emotion classification tasks. The results demonstrate that while hybrid architectures combining local spatial-temporal feature extraction with global attention mechanisms show promise, their performance remains inconsistent across tasks and datasets, with validation accuracies ranging from 37.33% to 66.67% .