This article proposes ActNet, a novel deep convolutional neural network architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address the problem of recognizing human actions from still images.
Abstract
Recognizing human actions from still images is a challenging task due to the absence of temporal information and the need to infer actions from subtle pose and contextual cues. In this article, we propose ActNet, a novel deep convolutional neural network (CNN) architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address this problem. ActNet integrates a Multi-Feature Network (MFNet) backbone for extracting rich features from multiple receptive fields, an Activity Multi-scale Block (AMB) for learning spatially diverse action patterns, and a Focus-Aware Recognition Module (FARM) that adaptively highlights the most informative regions of the image. We evaluate ActNet on the Stanford 40 Actions and PASCAL Visual Object Classes (VOC) 2012 datasets and show that it outperforms several state-of-the-art CNN and transformer-based models, achieving superior accuracy, precision, recall, and F1-score. Extensive ablation studies confirm the effectiveness of both AMB and FARM components. ActNet demonstrates robust generalization to a wide range of human actions, making it a strong candidate for still-image-based action recognition tasks in practical applications.
A novel action recognition method, named MICA-Net, which combines data from multiple sensors to improve the efficiency of the HAR model, and a new compact version of a wrist-worn sensor device with Wi-Fi connectivity to an edge device, enhancing usability in human-machine interaction applications.
Trung-Hieu Le, Thai-Khanh Nguyen, T. Tran et al.· ACM Transactions on Multimed...· 0 citations
This work proposes MSCALNet, a Multi-Scale Convolutional Attention LSTM Network, a Multi-Scale Convolutional Attention LSTM Network that employs a multi-branch differential encoding strategy to fuse heterogeneous sensor information, and efficiently models multi-timescale dynamics through a dilated convolutional pyramid...
Zi-Bo Wang, Runyang Lyu, Bin Xiao· International Conference on...· 0 citations
Advances in artificial intelligence have made hand gesture recognition an important human–computer interaction modality. Graph convolutional networks (GCNs) are widely used for skeleton-based hand gesture recognition, yet their performance can be limited by weak semantic topology modeling, underused feature channels, a...
Xiaowei Han, Ting-Shan Yan, Yunjing Lu et al.· Electronics· 0 citations
Weakly supervised video anomaly detection remains a challenging problem, primarily due to the scarcity of abnormal training samples and the lack of diverse feature representations, which hamper the learning of discriminative models. To address these issues, we introduce a novel weakly supervised cross-domain framework...
Mao-Wen Zhou, Erma Rahayu Mohd Faizal Abdullah, Aznul Qalid Md Sabri et al.· PLoS ONE· 0 citations
A Multi-Scale Convolutional Neural Network integrated with a Long Short-Term Memory network for sports video action recognition is proposed, demonstrating the potential of the proposed framework to combine spatial and temporal information for sports video action recognition.
Hai-Ming Yang, Hafiz Mohd Sarim, Xiao-Juan Ma et al.· Journal of Visualized Experi...· 0 citations
Experimental results on three public datasets demonstrate that MDSTA-Net achieves favorable recognition performance compared with several representative methods, indicating that multi-scale spatial feature enhancement and key spatiotemporal information aggregation can effectively improve action recognition performance.