Skip to content
Review Open access

Computer Vision Models for Human Activity Recognition: A Literature Review

Jun 2026 · Algorithms · 0 citations · 43 references

Abstract

Human Activity Recognition (HAR) is the automated process of identifying human actions using sensor data or video, which is widely used in healthcare, smart environments, and surveillance. Although HAR based on computer vision has advanced rapidly, existing reviews do not adequately address the recent shift toward hybrid deep-learning architectures or provide a structured comparison of the trade-offs relevant to real-world deployment. This literature review addresses that gap through a PRISMA-guided analysis of articles published between 2021 and 2025 and retrieved from four major databases. The review develops a reproducible taxonomy of nine architectural families and applies a multidimensional evaluation framework covering classification accuracy, computational efficiency for edge deployment, environmental generalization, and fine-grained activity recognition. The findings show that hybrid architectures are the dominant design strategy, while attention-based and graph-based models play important specialized roles depending on temporal complexity, privacy requirements, and deployment constraints, with the literature concentrated mainly in healthcare and security applications.

Read PDF

Similar papers

Conference Jul 2026

A Comprehensive Review of Human Activity Recognition Methods: Trends, Challenges, and Future Directions

Human Activity Recognition (HAR) is a fast-growing research area that focuses on identifying human actions using data collected from sensors and vision-based devices. It plays an important role in applications like health monitoring, smart homes, surveillance, sports analysis, and human-computer interaction. In recent years, several methods have been developed to improve the performance of HAR systems using machine learning, deep learning, and hybrid models. This paper presents a detailed review of different methods used in HAR. The study is divided into three main categories: vision-based methods, sensor-based methods, and hybrid approaches that combine both types. Each method is discussed with examples from recent research, along with their advantages and limitations. A comparison is also provided in the form of a table to highlight the performance and challenges of each approach. Although HAR systems have achieved good results in controlled environments, several challenges still remain. These include poor generalization to new users or unknown environments, difficulty in recognizing complex or overlapping activities, dependence on large datasets, and lack of real-time performance. This paper also discusses these research gaps based on recent findings. The future of HAR depends on building more accurate, reliable, and real-time systems that can adapt to different situations. The paper concludes by suggesting possible directions for future work, such as the development of lightweight models, use of standard datasets, better handling of real-time data, and making models more interpretable.

Satveer Kaur, Navneet Kaur Sandhu, Nitika Goyal · 0 citations
Review Open access Aug 2026

AI-based vision techniques for human activity recognition in surveillance videos

The article compares the performance of traditional machine learning techniques with recent deep learning architectures such as CNNs, RNNs, TCNs, and Transformers, based on accuracy, computational cost, and suitability for real-world disorderly plotting.

Disha Deotale, M. Verma, P. Suresh et al. · 0 citations
Open access Jul 2026

Sensor-Modality-Aware Human Activity Recognition with the Convolutional Tsetlin Machine: Interpretable and Resource-Efficient Neuro-Symbolic Learning

This work investigates the Convolutional Tsetlin Machine for multimodal HAR using only the raw inertial signals of the UCI-HAR dataset, rather than its pre-computed 561-feature representation, to address predictive performance, interpretability and suitability for embedded and mobile platforms.

O. Tarasyuk, A. Gorbenko, O. Gordieiev et al. · 0 citations
Preprint Jul 2026

RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment

This work introduces RAG-HAR+, a retrieval-first and cost-optimized extension that strengthens retrieval while reducing dependence on LLM-based inference, and extends the RAG-HAR mobile prototype to demonstrate the practical feasibility of retrieval-first, LLM-assisted HAR in mobile sensing scenarios.

Hansi Karunarathna, Nirhoshan Sivaroopan, Chamara Madarasingha et al. · 0 citations
Review Open access Jul 2026

WiFi-Based Human Activity Recognition and Fall Detection with Taxonomy, Benchmarks, and Future Directions: A Narrative Review

This review presents a comprehensive taxonomy of post-2017 architectures, a comparative synthesis of laboratory versus deployment performance, and a critical analysis of key implementation challenges, and outlines key recommendations for developing adaptive, location-independent models.

Kok Chung Chua, Kai Liang Lew, Chean Khim Toa et al. · 0 citations
Review Open access Aug 2026

Real-Time Vision-Based Fall Detection Systems for the Elderly: A Systematic Review

The findings show that CNN-based architectures dominate algorithm choice, edge devices dominate deployment platforms, and optimization remains central to real-time inference on constrained hardware.

Mahammad Nabizade, Réda Yahiaoui, Isabelle Lajoie et al. · 0 citations