Aug 2026· International Conference on Digital Image Processing· Vol 14351, pp. 143512H - 143512H-15· 0 citations· 16 references
Engineering
TL;DR
Tests show that this real-time intelligent monitoring, safety early warning, and analysis framework can quickly and accurately detect personnel, identify various actions, issue safety early warnings in a timely manner, and complete intelligent scene analysis, meeting the actual usage needs of kindergartens.
Abstract
Children's safety is the primary and core issue that kindergartens face. Traditional video surveillance is the main way for kindergartens to ensure safety, but this method highly relies on manual monitoring. It is not only inefficient but also prone to overlooking risks due to human negligence, and it is even more difficult to achieve real-time risk early warning. To solve these practical problems, we propose a real-time intelligent monitoring, safety early warning, and analysis framework based on a multi-modal large model. This framework integrates computer vision technology and large vision-language models, enabling both multi-dimensional scene perception and intelligent safety early warning and scene analysis. Specifically, for the needs of personnel identification and tracking, we fine-tuned the YOLOv11 model with a dedicated dataset to achieve high-precision real-time personnel detection; paired it with the ByteTrack algorithm to complete multi-target tracking, and then used the fine-tuned S3D network to identify children's dangerous actions or abnormal behaviors-grade these behaviors according to their danger levels and trigger corresponding preliminary early warnings. Then,we transmit these early warning results and scene images to the Qwen3-VL model for scene-level risk reasoning and finally generate a complete safety analysis report. We conducted tests in real kindergarten scenarios, and the results show that this framework can quickly and accurately detect personnel, identify various actions, issue safety early warnings in a timely manner, and complete intelligent scene analysis, meeting the actual usage needs of kindergartens.
Martial arts routines are highly dynamic, multi-pose, and fast-paced, posing significant challenges to automated recognition and scoring. Such complex spatiotemporal characteristics are also representative of intelligent sensing tasks in advanced electromagnetic-aware environments, where robust human motion perception is essential for integrated visual and wireless monitoring systems. This paper proposes a joint framework combining YOLOv8 and time-optimized OpenPose to mitigate pose estimation jitter and detection inaccuracies caused by rapid motion and occlusion in human motion analysis. YOLOv8 first performs high-precision human target detection using bounding boxes to define regions of interest, after which the cropped images are processed by OpenPose to extract 17 keypoints, with IoU matching ensuring cross-frame identity association for temporal consistency. The keypoint sequence is modeled as a multi-channel time series and refined through a bidirectional LSTM network to predict smooth pose trajectories. The optimized keypoints are further used to calculate joint angles and movement velocities, which are integrated with dynamic thresholds for motion segmentation. DTW-based alignment and similarity matching with a standard motion library are subsequently employed to evaluate posture accuracy, motion amplitude, and rhythmic consistency, producing a comprehensive scoring result. Experimental results demonstrate average mAP@0.5 values of 0.863–0.942 (FPS 115.2–124.3), average PCK values of 88.7%–90.3% under occlusion (average MPJPE 48.2–51.6 pixels), and Pearson correlation coefficients of 0.925–0.941 for scoring consistency across diverse martial arts routines. The proposed framework provides an effective solution for intelligent motion analysis and offers technical reference for multimodal perception and dynamic scene understanding in advanced electromagnetic sensing applications.
D. Zhao, Y. Q. Ma· Advanced Electromagnetics· 0 citations
Due to the limitations of visual surveillance systems, such as line-of-sight, obtrusiveness, and high power requirements, audio-based surveillance has gained significant traction in security and forensic applications. Among these, footstep-based audio analysis has emerged as a promising and non-intrusive approach for monitoring and threat detection. This paper introduces EWFootstep 1.0, a novel dataset comprising recordings from 176 subjects collected under real-world environmental conditions, distinguishing footstep acoustic signatures of single and multiple individuals across forests, roads, and indoor settings. To validate the dataset, we perform time and frequency domain analyses, and implement machine learning (ML) and convolutional neural network (CNN) based baseline models. Feature separability is visualized using t-SNE and quantified using Davies-Bouldin Index. Our work bridges the gap in footstep-based security by providing a comprehensive dataset for machine learning applications in surveillance and crime scene investigation.
Anshuman Agrahri, C. Maurya, Ravi Shekhar Tiwari et al.· Journal on Audio, Speech, an...· 0 citations
This study suggests an enhanced safety helmet detection method based on YOLOv10 to solve the low detection accuracy of current algorithms for small objects and complicated settings in different situations.
Iqra Aziza Khatoon, Dr. Safia Khanam· International Journal of Dat...· 0 citations