A proactive, real-time computer vision system designed to detect potentially suspicious behavior around parked vehicles, with a specific focus on unauthorized proximity and loitering is proposed, making it a strong candidate for practical urban vehicle monitoring, subject to further large-scale validation across diverse environments.
Abstract
Vehicle theft and vandalism remain significant urban security challenges commonly addressed through reactive, post-incident forensic measures. This paper proposes a proactive, real-time computer vision system designed to detect potentially suspicious behavior around parked vehicles, with a specific focus on unauthorized proximity and loitering. The proposed architecture integrates state-of-the-art object detection using YOLOv11 (You Only Look Once version 11), multi-object tracking via a lightweight custom association tracker inspired by the ByteTrack/StrongSORT/OC-SORT paradigm, and monocular depth estimation based on the Intel DPT-Large framework.A key contribution is the identification and mitigation of the Perspective Challenge: the two-dimensional (2D) scale ambiguity that causes distant background pedestrians to appear falsely proximate to foreground vehicles in monocular camera feeds. To address this, three spatial analysis strategies are implemented and evaluated: (A) fixed Euclidean thresholding, (B) adaptive perspective thresholding, and (C) three-dimensional (3D) depth injection. Experimental results on real-world urban surveillance footage (27,000 annotated frames across two datasets) demonstrate that Strategy C achieves the highest precision (0.95) with an F1-score of 0.92, while Strategy B provides the best balance between accuracy (precision 0.88, recall 0.91, F1 0.89) and computational efficiency (32.7 frames per second, FPS). Compared to naive 2D thresholding (Strategy A), Strategy B reduces false alarms by approximately 80%, while Strategy C further improves precision to 0.95 through depth-plane verification. The system maintains real-time performance exceeding 30 FPS under Strategy B, making it a strong candidate for practical urban vehicle monitoring, subject to further large-scale validation across diverse environments.
This study proposes an Advanced Surveillance Framework that makes use of YOLOv10, a next-generation real-time object detection algorithm that greatly outperforms conventional single-sensor approaches in precision, recall, and real-time responsiveness.
Sadiya Begum, Lubna Nausheen, Ruqiya Fatima· International Journal of Eng...· 0 citations
The advancement of military internet of things (IoT) surveillance demands real-time casualty detection across distributed camera networks under diverse environmental conditions. Traditional surveillance systems suffer from high latency, bandwidth inefficiency and unreliable cross-camera identity tracking, indicating the need for advanced detection and tracking systems. This study presents an edge-based multicamera casualty detection and tracking system for military IoT networks. The proposed framework, called CCTV-TrackNet, integrates lightweight AI on distributed camera nodes built on Raspberry Pi Zero2W hardware with high-resolution cameras and sensors such as GPS and IMU. At each node, YOLOv12n performs real-time person and casualty detection, DeepSORT maintains short-term continuity, and OSNet-based Re-ID extracts appearance embeddings for cross-camera association. A central server aggregates metadata for global identification, visualizes trajectories, and generates real-time alerts. Experimental results show 90.77% detection accuracy, 37% and 36% reduction in false cases, 30.3 FPS edge performance, and 84.20% cross-camera ID consistency.
Sium Bin Noor, M. A. Dini, Jaemin Lee et al.· International Conference on...· 0 citations
Rapid urbanization has increased the need for surveillance systems that can monitor multiple public safety risks at the same time. Traditional systems often use separate solutions for facial recognition, vehicle identification, fire detection, and behavioral analysis, resulting in fragmented infrastructure and multiple interfaces for operators to manage. This paper presents City Sentinel, a unified AI-based surveillance framework that integrates six detection capabilities into one scalable platform: facial recognition, automatic number plate recognition (ANPR), fire and smoke detection, weapon and knife detection, violence detection, and road accident detection. The system combines a Next.js operator dashboard, FastAPI backend, cloud-based PostgreSQL event storage, InsightFace and YOLOv8 vision models, and EasyOCR for plate recognition. Camera streams are processed through dedicated inference workers using RTSP. On a workstation equipped with an NVIDIA RTX 3060 GPU, the system achieves a median end-to-end latency of 743 ms and supports four concurrent RTSP streams within a two-second latency limit. It achieves a 91.2% face-match rate, 85.7% plate-reading accuracy, and mAP@0.5 scores of 0.846 to 0.889 across the fire, knife, and weapon detection modules. In user-acceptance testing, operators could enroll a new identity in under one minute and identify a flagged person from live footage in an average of 12 seconds. The results demonstrate that a modular, open-source, multi-model architecture can provide broad surveillance coverage, cloud-based auditability, and flexibility for adding new detection capabilities while maintaining practical real-time performance.
Hanan Syed Shabir, Noor Fatima, Safia Baloch et al.· 0 citations
AI-based visual perception systems are increasingly deployed in infrastructure surveillance, including roadside monitoring units, highway cameras, and smart-city pedestrian management systems. The security vulnerability of these systems to physical adversarial attacks poses a direct threat to the reliable operation of transportation infrastructure. We propose AdvSerial, a dynamic 2D--3D joint optimization framework for generating continuous high-angle physical adversarial patches against pedestrian detectors in infrastructure-based scenarios. We UV-map a boundary-aware quilted texture onto 3D garments, combine 2D digital attacks with 3D sparse- and continuous-frame rendering, and explicitly suppress person-specific semantic features while enforcing temporal continuity. A Feature Smooth Quilting strategy reduces visible patch boundaries and bounds cross-seam feature discontinuities. A serial-frame loss encourages long uninterrupted sequences of detection failures. In physical world experiments, AdvSerial achieves a 74.8% attack success rate on YOLO-v5 and degrades mean detection confidence from 84.30% to 39.38%. Experiments spanning eight detectors with different architectures demonstrate strong transferability. Notably, it achieves an $89.71%$ attack success rate on YOLO-v2 and resists both patch-detection defenses (NapGuard) and 3D-temporal perception (Sparse4D-v3). The results reveal persistent, temporally consistent failure modes under high-angle surveillance, and motivate the design of motion-aware and 3D-aware defenses for security-critical infrastructure deployments.
Yuanhao Huang, Yilong Ren, Jinlei Wang et al.· 1 citation
Reliable traffic sign detection is a prerequisite for the global deployment of autonomous driving systems, where regulatory compliance and road safety depend on perceiving signs correctly across regions, ranges, and weather conditions. Despite recent progress, vision-based methods continue to face three fundamental limitations: poor cross-regional generalization due to high diversity across countries, degraded performance on small-object detection at long ranges (traffic signs occupy as little as $10{\times}10$ pixels at 200m), and fragile temporal tracking under the strongly non-linear perspective distortion that occurs as a vehicle approaches a sign. In this paper, we address the problem of robust, long-range, region-agnostic traffic sign perception by combining camera and Light Detection and Ranging (LiDAR) sensing. We present a multi-modal detection framework whose Intensity-Aware Deformable Fusion module aligns retro-reflective LiDAR cues with camera features, anchoring detection on geometric invariants rather than region-specific visual appearance. We further introduce a dual motion-model tracker that explicitly accounts for non-linear perspective transformations during vehicle approach, substantially improving temporal consistency over linear motion assumptions. Additionally, we develop a semantic attribute classification pipeline that estimates occlusion level, readability, sign embeddedness, and road relevance, providing actionable context to downstream planning. Extensive evaluation on our dataset, spanning 60+ countries and 2,500+ hours of driving data, shows that the proposed pipeline achieves an Object Miss Ratio (OMR) of 0.49% across 221,068 evaluation sequences, demonstrating globally generalizable traffic sign perception in commercial-grade autonomous driving systems.
Meda Lazar, S. Sridhar, Shashwata Gupta et al.· 0 citations
Intelligent transportation systems, traffic surveillance and smart city monitoring require accurate vehicle detection and tracking. Traditional methods of monitoring are usually based on either manual monitoring or GPS-based monitoring which may have issues with signal dependency and lack of scalability. This paper suggests a deep learning-based car surveillance system, which combines the YOLOv8 object detector and a multi-object tracking system to perform automated automobile detection and tracking. A prepared set of custom traffic image data (about 1,700 images) was divided into a training (80%) and validation (20%) sample and trained on the YOLOv8-Nano model to detect vehicles. Images had been resized to 640 × 640 resolution during training with 50 epochs of a batch size of 16 with transfer learning on pretrained weights. The trained detector had the accuracy of 0.81 with the recall of 0.74 and the mean Average Precision (mAP-0.5) of 0.79 on vehicle detection assignments. The detecting unit was further combined with tracking structure to retain vehicle identities between sequential frames to provide the capability of regular tracking of multi-object in traffic environment. The test outcomes show that the suggested system can work at about 40 frames per seconds (FPS) with the evaluation dataset and still retain a good tracking precision of about 0.76. Bounding boxes, tracking IDs and performance graphs are some of the visualization results, which support the efficiency of the methodology. The given framework can also be used to track the location of military vehicles in surveillance domains during the situations when the convoy movements or tactical vehicle location can be monitored automatically in GPS-denied or irregular conditions.
S. Karunya, B. Praveen, VJ Sara Belina· ITEGAM- Journal of Engineeri...· 0 citations