A Multimodal Gesture Recognition Framework Using sEMG and ACC Signals with Cross-Temporal Attention and Mutual Information Regularization
Abstract
Surface electromyography (sEMG), which reflects temporal muscle activity, has been widely applied in gesture recognition and human-computer interaction. However, single-modal sEMG signals are susceptible to factors such as electrode displacement, inter-subject variability, and noise, which limit recognition performance. To improve system robustness, this paper proposes a multimodal gesture recognition method based on sEMG and acceleration (ACC) signals. First, temporal convolutional encoders are employed to extract temporal features from the sEMG and ACC signals, respectively. Then, a Cross-Temporal Attention mechanism is introduced to fuse features from different modalities, thereby modeling the temporal correlations between multimodal signals. In addition, Mutual Information Neural Estimation (MINE) is adopted to constrain the mutual information between features of different modalities, so as to enhance the consistency and discriminative capability of the fused features. To verify the effectiveness of the proposed method, a multimodal gesture dataset containing 11 subjects, 10 gestures, and 4 different postures was constructed. Experiments were conducted using five-fold cross-validation within each subject. Experimental results show that, compared with single-modal methods and simple fusion methods, the proposed method achieves superior performance in both recognition accuracy and Macro-F1 score, with an average recognition accuracy of 92.41%. These results demonstrate that the proposed method can effectively improve multimodal gesture recognition performance and has strong application potential.