Skip to content
Open access

Action Recognition in Sports Videos Using Multiscale Convolutional Networks with Long Short-Term Memory (LSTM)-Based Temporal Modeling.

Sep 2026 · Journal of Visualized Experiments · Vol 235 · 0 citations
Medicine

TL;DR

A Multi-Scale Convolutional Neural Network integrated with a Long Short-Term Memory network for sports video action recognition is proposed, demonstrating the potential of the proposed framework to combine spatial and temporal information for sports video action recognition.

Abstract

Action recognition in sports videos remains challenging because of complex motion dynamics, occlusion, and high intra-class variability. Although existing deep learning approaches, including CNN-BiLSTM and transfer learning-based models, have demonstrated effectiveness in human activity recognition, their performance may be limited in sports scenarios with rapid, diverse movements. Many existing methods rely on single-scale convolutional filters, which may not effectively capture both fine-grained and coarse motion characteristics simultaneously. To address this limitation, this study proposes a Multi-Scale Convolutional Neural Network (MSCNN) integrated with a Long Short-Term Memory (LSTM) network for sports video action recognition. The MSCNN extracts spatial representations at multiple receptive fields through parallel convolutional kernels, enabling the learning of both detailed and contextual motion features. These features are subsequently processed by the LSTM to capture temporal dependencies and motion continuity across consecutive frames. Experimental evaluation was conducted on the UCF11, UCF Sports, and JHMDB benchmark datasets. The proposed MSCNN-LSTM model achieved classification accuracies of 98.3%, 95.4%, and 81.7%, respectively, outperforming the comparative approaches evaluated in this study. An ablation study further demonstrated the contribution of multi-scale feature extraction and temporal modeling to overall performance. These findings demonstrate the potential of the proposed framework to combine spatial and temporal information for sports video action recognition.

Read PDF

Similar papers

Open access Aug 2026

Action Recognition Method Based on Multi-Scale Dilated Feature Fusion and Decoupled Spatiotemporal Attention Pooling

Experimental results on three public datasets demonstrate that MDSTA-Net achieves favorable recognition performance compared with several representative methods, indicating that multi-scale spatial feature enhancement and key spatiotemporal information aggregation can effectively improve action recognition performance.

Han-Bo Zhang, Jing Huang · 0 citations
Open access Sep 2026

ActNet: focus-aware multi-scale CNN for human activity recognition from images

This article proposes ActNet, a novel deep convolutional neural network architecture that combines multi-scale feature learning with a focus-aware attention mechanism to address the problem of recognizing human actions from still images.

Şafak Kılıç · 0 citations
Conference Aug 2026

ActionLMM: captioning long-video actions with memory-augmented VLMs

This work introduces ActionLMM, a memory-augmented vision-language model for long-video action summarization that aligns visual and motion modalities through joint representation learning and leverages a novel dual-memory mechanism to retain both local motion details and global temporal structure.

Rui-Rui Li, Dari Abdullah Alrwoaily, Turgut Sofuyev et al. · 0 citations
Preprint Sep 2026

Few-Shot Video Recognition via Hierarchical Metric Learning

Few-shot action recognition (FSAR) aims to recognize unseen action categories with only a small number of annotated video samples. Recent works typically apply single-prototype supervision at the network output and fail to sufficiently exploit rich cross-frame global spatial information in videos. Even existing multi-l...

Jia-Xin Zhang, Hao-Ran Gao, Xi-Zhan Gao et al. · 0 citations
Open access 2026

Cross-Modal Temporal Alignment for Action Grounding in Videos

A cross-modal temporal alignment framework that combines a multi-scale temporal convolutional encoder with capsule-based dynamic routing, jointly optimizing temporal boundary prediction, cross-modal semantic alignment, and capsule diversity is proposed.

Gengtian Shi, Chen-Hao Wu, Shao-Fei Wang et al. · 0 citations
Open access Aug 2026

AN EFFICIENT MULTIMODAL FRAMEWORK FOR HUMAN ACTION RECOGNITION USING TCN AND BILSTM WITH LATE FUSION

In recent years, researchers have become interested in human action recognition because of the broad spectrum of its applications in surveillance, healthcare monitoring, and human-computer interaction. A highly robust, efficient, and innovative multimodal system for recognizing human actions is proposed in this study b...

Vijay Singh Rana, Ankush Joshi, K. Verma · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.