In complex environments, sign language recognition (SLR) is easily affected by background clutter, motion blur, and rapid movement. These factors can obscure subtle gesture patterns and weaken the discriminative spatiotemporal features required for recognition. To address these challenges, we propose SST-Net, a spatiot...
Han-Bo Zhang, Jie Miao, Qiu-Hong Tian et al.· Electronics· 0 citations
Video action recognition requires the joint modeling of spatial appearance information and temporal dynamics. However, existing efficient action recognition methods based on two-dimensional convolution still have limitations in representing multi-scale spatial cues and aggregating key spatiotemporal information. To add...
Human pose estimation in factory surveillance is challenged by scale variation, limb occlusion, complex backgrounds, and spatial detail loss during upsampling. This paper proposes MFA-Pose, an improved YOLO11s-Pose framework that enhances contextual representation and cross-scale feature reconstruction. The Multi-Scale...