Skip to content
Open access

AN EFFICIENT MULTIMODAL FRAMEWORK FOR HUMAN ACTION RECOGNITION USING TCN AND BILSTM WITH LATE FUSION

Aug 2026 · JOURNAL OF MECHANICS OF CONTINUA AND MATHEMATICAL SCIENCES · 0 citations · 18 references

Abstract

In recent years, researchers have become interested in human action recognition because of the broad spectrum of its applications in surveillance, healthcare monitoring, and human-computer interaction. A highly robust, efficient, and innovative multimodal system for recognizing human actions is proposed in this study based on the Florence 3D Action dataset. The proposed system leverages the RGB, Depth, and Skeleton multi-modalities. For the case of RGB and Depth video sequences, spatial information is extracted by a pre-trained ResNet50 model, and then the extracted high-level visual features are fed to a Temporal Convolutional Network (TCN) to capture temporal patterns at a large scale. Color and depth video sequences are processed in parallel. Temporal information is also captured in the skeleton data, which is fed to a Bi-directional Long Short-Term Memory (BiLSTM) model after features are extracted in the form of joint angles, velocities, and accelerations, which describe the motion of human actions. For the case of the three modalities, the results are fused at the decision level using a weighted sum for the final prediction. The results show that the proposed system has an impressive recognition accuracy of 96.8% on the Florence 3D dataset, with stable convergence, low validation loss and overfitting, and a strong generalization capability, which meets the requirements for use in real-time, resource-constrained systems.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.