Aug 2026· Neural Networks· Vol 205 Pt B, pp.
109517
· 0 citations· 53 references
Medicine
TL;DR
This work proposes MEMC (Masked Modeling with Efficient and Minimal Contrastive Learning), a novel framework that adopts an efficient sequential cascade strategy based on layer-grafted pretraining and introduces two CL enhancements to improve the discriminative capability of the learned representations.
Abstract
Unsupervised 3D skeleton-based action recognition offers advantages in terms of robustness and computational efficiency. However, prevailing paradigms, such as masked skeleton modeling (MSM) and contrastive learning (CL), have inherent limitations: MSM often learns representations with limited discriminative capability, whereas CL may fail to adequately capture fine-grained structural information. Furthermore, integrating the two paradigms through multi-task learning (MTL) can lead to gradient conflicts that limit performance. To address these issues, we propose MEMC (Masked Modeling with Efficient and Minimal Contrastive Learning), a novel framework that adopts an efficient sequential cascade strategy based on layer-grafted pretraining. MEMC first uses MSM to learn low-level representations and then refines high-level representations through CL, thereby avoiding the gradient conflicts associated with MTL. To improve the discriminative capability of the learned representations, particularly for subtle actions, we introduce two CL enhancements. First, Topology Distance-Aware Chain Motion Modeling incorporates skeletal topology priors to capture discriminative motion patterns along skeletal chains. Second, Frequency Band-Aware Contrastive Learning with Frequency Band Pooling (FBP) separates and integrates high-frequency details with low-frequency global context. Extensive experiments on three benchmark datasets demonstrate the effectiveness of MEMC and show its superior performance compared with state-of-the-art methods.
This work presents a novel self-supervised architecture centered on graph prototype learning that sets a new state-of-the-art on the ARMM dataset with an accuracy of 95.70%, substantiating the efficacy and transferability of prototype-guided self-supervised learning for skeleton-based action representation.
Zhijie Xu, Hongwei Chen, Xia Li· International Journal of Mac...· 0 citations
The Spatial-Temporal Graph Convolutional Network (ST-GCN) has achieved remarkable success in skeleton-based action recognition. However, it exhibits a systematic limitation in fine-grained scenarios, frequently misclassifying subtle, locally driven actions as common whole-body dominant patterns, resulting in persistent semantic confusion. To overcome these shortcomings, we introduce the Adaptive Joint Feature Enhancement (AJFE) module, which learns joint-specific importance weights for improved spatial aggregation, and the Multi-Scale Temporal Sensitivity Enhancement (MTSE) module, which captures multi-scale dynamics via parallel convolutional branches with adaptive fusion. Extensive experiments on NTU RGB+D and Kinetics demonstrate that our approach delivers competitive overall performance (84.5%/88.8% on Cross-Subject/Cross-View protocols) while substantially enhancing semantic consistency in challenging fine-grained cases. These lightweight enhancements offer a practical and effective solution for reliable skeleton-based action recognition in real-world applications.
Xin Han, Lin Zhang, Xin Chen et al.· International Conference on...· 0 citations
Efficient Point Masked Autoencoders (EP-MAE), a new framework designed to significantly reduce the training cost of 3D self-supervised pre-training while maintaining strong representation quality, and provides a scalable and effective foundation for future 3D neural network models is presented.
Jian Zhu, Jiale Zhao, Cheng Lin et al.· Neural Networks· 0 citations
Fine-grained human action recognition (FHAR) must distinguish visually similar actions that differ mainly in body configuration, timing, or local appearance. RGB representations retain visual context but often suppress joint-level geometry, whereas skeleton representations encode kinematics but discard dense spatial detail. We introduce FineX, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology. Pairwise cross-attention enables symmetric, stream-preserving information exchange, followed by a streamwise latent sparse Mixture-of-Experts that routes each representation to a content-dependent subset of shared experts, regularized by a load-balancing objective. FineX achieves state-of-the-art results on Gym99, Gym288, and Diving48. On the long-tailed Gym288, it raises mean class accuracy from 68.6% to 76.2% (+7.6 points) without textual supervision or large-scale vision-language pre-training, demonstrating the benefit of structured visual-pose-graph fusion and conditional expert refinement for FHAR.
Imtiaz ul Hassan, Tasweer Ahmad, Nikolaos Bessis et al.· 0 citations