Adaptive Convolutional Neural Network-Enhanced Scale-Fusion Network for Human Activity Recognition Using Wearable Sensors
Abstract
Deep learning has achieved notable success in sensor-based human activity recognition (SHAR), yet wearable inertial signals contain activity patterns at different temporal scales while practical deployment imposes strict computational constraints. This paper proposes a three-stage Adaptive CNN-Enhanced Scale-fusion Network (ACESNet) for six-channel accelerometer–gyroscope activity recognition. The Adaptive Kernel Encoder dynamically combines temporal kernels of different sizes, while the CNN-Enhanced Scale-fusion design combines multi-scale temporal mixing, explicit channel mixing and a parallel local CNN path without materializing a global attention matrix. The main benchmark uses a fixed stratified 70/15/15 window-level partition (split seed 42), and all methods in the main comparison are repeated over five training seeds. On REALWORLD, MotionSense, UCI-HAR and SHL, ACESNet achieves accuracies of 95.37 ± 0.09%, 99.33 ± 0.12%, 98.01 ± 0.23% and 93.49 ± 0.10%, with Macro-F1 scores of 95.59 ± 0.09%, 99.12 ± 0.15%, 98.16 ± 0.21% and 94.03 ± 0.08%, respectively. A separate official subject-independent UCI-HAR evaluation yields 94.01 ± 0.84% accuracy and 94.11 ± 0.88% Macro-F1. The adaptive-kernel selector has a mean normalized entropy of 0.9114, providing evidence against severe single-branch collapse rather than strong kernel selectivity. On a Raspberry Pi 5 (Raspberry Pi Ltd., Cambridge, UK), ACESNet requires 8.731 MFLOPs and 1.606 ± 0.014 ms per window in FP32 ONNX Runtime 1.18.1. These results support a favorable performance–efficiency trade-off relative to the selected baselines under the stated protocol.