Adaptive joint weighting and multiscale temporal modeling for fine-grained skeleton action recognition
Abstract
The Spatial-Temporal Graph Convolutional Network (ST-GCN) has achieved remarkable success in skeleton-based action recognition. However, it exhibits a systematic limitation in fine-grained scenarios, frequently misclassifying subtle, locally driven actions as common whole-body dominant patterns, resulting in persistent semantic confusion. To overcome these shortcomings, we introduce the Adaptive Joint Feature Enhancement (AJFE) module, which learns joint-specific importance weights for improved spatial aggregation, and the Multi-Scale Temporal Sensitivity Enhancement (MTSE) module, which captures multi-scale dynamics via parallel convolutional branches with adaptive fusion. Extensive experiments on NTU RGB+D and Kinetics demonstrate that our approach delivers competitive overall performance (84.5%/88.8% on Cross-Subject/Cross-View protocols) while substantially enhancing semantic consistency in challenging fine-grained cases. These lightweight enhancements offer a practical and effective solution for reliable skeleton-based action recognition in real-world applications.