Computer vision-based human motion recognition: technological evolution, key challenges, and cutting-edge trends
Abstract
Human motion recognition is an important research direction in computer vision, aiming to automatically analyze human pose changes and action semantics from images or videos captured by visual sensors. This paper systematically reviews the evolution of computer vision-based human motion recognition technology, from traditional handcrafted feature methods to convolutional neural networks and recurrent neural networks in the deep learning era, and then to the recently emerging Transformer architectures and vision-language models. The core ideas and technical characteristics of various approaches are comprehensively analyzed. On this basis, key challenges in current research are deeply explored, including the complexity of spatiotemporal feature extraction, robustness issues with occlusion and viewpoint changes, difficulties in fine-grained action and few-shot learning, and the trade-off between multimodal fusion and computational efficiency. Finally, future development trends are discussed, pointing out that cutting-edge directions such as 3D human pose estimation, fisheye lens adaptation, micro-action detection, and attribute-aware generation will drive the field toward higher precision, stronger generalization capabilities, and broader application scenarios. This paper aims to provide systematic technical reference and forward-looking insights for researchers in human motion analysis and intelligent human-computer interaction.