Skip to content
Open access

A Multi-Scale Dual-Head YOLOv5 Framework for Hand Gesture Recognition via Spatial Relationship Modeling

Sep 2026 · Information · 0 citations · 8 references

Abstract

Hand gesture recognition plays an important role in computer vision with broad applications in human–computer interaction, intelligent perception, and contactless interaction systems. However, existing detection-based methods mainly rely on hand appearance features, which are often insufficient for distinguishing gesture categories with similar local hand shapes but different hand–head spatial configurations. To address this issue, a multi-scale dual-head hand gesture detection framework based on YOLOv5s is proposed. The framework introduces an auxiliary detection branch to explicitly detect hand and head regions, and the corresponding multi-scale feature representations are fused to incorporate hand–head spatial contextual information into gesture recognition. Furthermore, Spatial Pyramid Pooling (SPP) and a Convolutional Block Attention Module (CBAM) are employed to enhance multi-scale contextual representation and feature discrimination before final gesture prediction. To evaluate the proposed framework, experiments were conducted on a HaGRID-based Dataset, a UAV-Gesture Dataset, and a self-collected real-world zero-shot dataset. Experimental results demonstrate that the proposed framework consistently improves detection performance over the YOLOv5s baseline while maintaining relatively lightweight model complexity and real-time inference capability. In addition, ablation studies verify the effectiveness of the proposed dual-head architecture, multi-scale feature fusion, SPP, and CBAM, while cross-dataset and zero-shot evaluations further demonstrate the applicability of the proposed framework across different gesture datasets and unseen real-world scenarios.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.