Skip to content

Automatic traffic surveillance system leveraging generative AI large language models

TL;DR

A generative AI-based traffic surveillance system leveraging large language models (LLMs) to enable timely and context-rich interpretation of traffic events, demonstrating the system's effectiveness in producing context-aware traffic scene descriptions, improving operational decision-making, and enhancing roadway safety.

Abstract

The application of automated traffic surveillance systems has become increasingly critical for improving traffic management and reducing reliance on manual monitoring. Traditional video surveillance methods are time-consuming, resource-intensive, and prone to human error, with operators frequently missing incidents due to fatigue and environmental factors. This study introduces a generative AI-based traffic surveillance system leveraging large language models (LLMs) to enable timely and context-rich interpretation of traffic events. We developed a custom-annotated dataset of 604 videos, capturing crashes, congestion, lane closures, and diverse weather and lighting conditions from Missouri Department of Transportation cameras and online sources. This approach fine-tunes the Qwen2.5VL-Instruct model using Low-Rank Adaptation (LoRA), temporal context enhancement, and multidimensional Rotary Position Embedding (mRoPE) for improved cross-modal fusion. Compared with the LLaVA-NeXT-Video baseline, the fine-tuned model achieves substantial gains in captioning performance (BLEU-4 = 0.4569, METEOR = 0.6187, CIDEr = 5.4612) and reduces average video description time from 35 seconds (manual) to 20 seconds (automated). An interactive web interface integrates real-time traffic monitoring with automatic incident detection and visualization, supporting faster emergency response and scalable deployment across traffic management centers. These results demonstrate the system's effectiveness in producing context-aware traffic scene descriptions, improving operational decision-making, and enhancing roadway safety.

View source

Similar papers

Conference Jul 2026

Query-Driven Intelligent Surveillance using Deep Learning for Activity Recognition and Video Summarization

Surveillance systems have experienced rapid growth which results in production of large video data streams. The monitoring process for this data becomes challenging because its volume exceeds human capacity and this situation creates potential for errors. Our research presents a hybrid intelligent surveillance system which conducts automatic video analysis through its two core operational components. The system employs two primary components to achieve its objectives. The SlowFast-based model enables users to track activities through their development across various time intervals. The system employs YOLO-based models to identify critical objects which include fire and weapons and road accidents through real-time monitoring. The system achieves improved stability through the implementation of a temporal debouncing method. The system uses multiple frame detection checks to improve detection accuracy which helps prevent false alarms. The system includes a module dedicated to video summarization which creates a summary from detected activities and visual changes. The system discards unneeded video content while retaining essential information through this process. The model uses a dataset that contains 4758 video clips which display various classification types. The system reaches 85% validation accuracy which demonstrates its ability to handle new data successfully. The system operates on devices with limited resources while providing an immediate alert system to inform users about essential incidents. The system delivers an easy-to-use and effective solution for intelligent video surveillance operations.

Abdul Haq Nalband, R. U, Shashwat Dodamani et al. · 0 citations
Review Open access Jul 2026

A Deep Neural Network Framework for Automated Crime Detection and Video Summarization

The exponential growth of closed-circuit television (CCTV) infrastructure has far outpaced the capacity of human operators to monitor footage in real time, leaving a critical gap between data acquisition and actionable situational awareness. This paper proposes a unified deep neural network framework that jointly performs automated crime (anomalous-event) detection and attention-guided video summarization from untrimmed surveillance streams. The framework couples a 3D convolutional feature extractor with a bidirectional long short-term memory (BiLSTM) temporal encoder and a self-attention module to model both short-range spatial cues and long-range temporal dependencies characteristic of criminal activities such as assault, robbery, arson, and shooting. The learned frame-level attention scores are re-used, without additional supervision, to drive a diversity-aware keyframe selection module that automatically compresses hours of footage into a compact, timestamped evidentiary summary whenever an anomalous segment is detected. The proposed system was evaluated on the UCF-Crime and DCSASS benchmark datasets for detection and on TVSum/SumMe-style protocols for summarization quality. Experimental results indicate that the framework attains a detection accuracy of 98.4%, an F1-score of 0.978, and an AUC of 0.991, outperforming C3D, two-stream CNN, and CNN-LSTM baselines by 3.4-16.0 percentage points, while the attached summarization module achieves an F-score of 0.71 with a video compression ratio exceeding 92% and end-to-end inference throughput of 41 frames per second on a single consumer-grade GPU. These results demonstrate that combining anomaly-aware attention with keyframe extraction in a single trainable pipeline yields both higher detection fidelity and immediately actionable, human-reviewable video summaries, making the framework well suited for real-time deployment in smart-city and public-safety surveillance infrastructures.

Kanika Singhal, Deepak Chandra Uprety, Amrita Bhatnagar et al. · 0 citations
Open access Jul 2026

ULSTM: Multi-Scale and Full-Level Temporal Consistency for Traffic Anomaly Detection

Experimental results demonstrate that the ULSTM framework significantly outperforms frame-independent generative models by suppressing high-frequency reconstruction noise, providing a robust solution for real-world smart city deployments.

Borja Pérez, Mario Resino, Jaime Godoy et al. · 0 citations
Conference Jul 2026

Traffc Signs Detection using YOLOv8: A Real-Time Deep Learning Approach for Autonomous Driving Systems

Traffic Sign Detection and Recognition (TSDR)— the automatic localisation and classification of traffic signs from onboard cameras—is a fundamental component of Advanced Driver Assistance Systems (ADAS) and autonomous vehicles: approximately 1.3 million road fatalities occur annually worldwide, with driver inattention to traffic signs ranked the second most common contributing factor after speeding. However, existing TSDR approaches exhibit notable limitations: two-stage and transformer-based detectors are too slow for real-time onboard inference, anchor-based single-stage detectors require extensive anchor tuning and degrade sharply on small and distant signs, and most published systems neglect deployment readiness and region-specific sign diversity such as Indian signage. To overcome these problems, we propose a complete real-time TSDR pipeline built on YOLOv8, combining an anchor-free prediction head, a C2F backbone, and a Path Aggregation Network (PANet) neck, complemented by a multi-task CNN (EfficientNet-v2 with Online Hard Example Mining) for Indian traffic signs and ONNX export for edge deployment. In this paper, the pipeline is trained and evaluated end-to-end, and its results are positioned against a structured survey of the 2025–2026 state of the art (YOLO-BS, CPB-YOLOv8, FEBG-YOLOv8s, CRS-NET, and an optimised YOLOv7). Experimental results show that the proposed pipeline achieves Precision: 90.6%, Recall: 88.9%, and mAP@0.5: 93.1% (validation) / 92.3% (test) after only 30 epochs on the primary dataset; on the CCTSDB2021 benchmark, YOLOv8 attains mAP@0.5: 97.1% on the test split, outperforming both YOLOv7 and YOLOv9, while the multi-task CNN achieves state-of-the-art F1 = 98.0% on the Indian traffic sign dataset. The exported ONNX model retains full accuracy at 128 FPS on GPU, demonstrating a deployment-ready real-time system.

Hussain Bhanpurawala, Sonali Ajankar · 0 citations
Open access Aug 2026

Amneen: An AI-Based System for Real-Time Tracking and Management of Crowds

With increasing crowd sizes nowadays, the risks of overcrowding, including injuries, accidents, and even fatalities, have become a great concern. Motivated by the need for safer public spaces, this work designed and developed Amneen, a crowd management system that uses Artificial Intelligence (AI) and Computer Vision (CV). The system provides authorities and event organizers with a real-time tool to track crowd density in public places and prevent dangerous situations before they occur. Amneen integrates two AI models: a head detection model using YOLO-11 and an overcrowding prediction model using Stochastic Gradient Descent Regressor (SGDRegressor). Using live video footage from installed cameras, the system estimates the number of people in a specific area, displays crowd statistics through an interactive dashboard, and sends early warnings when the situation worsens. Additionally, by analyzing historical data patterns, the system predicts congestion before it occurs. The detection model demonstrated strong performance in real-time, processing an image in 6.5 ms with a precision of 93.95%, a recall of 90.91%, an F1-score of 92.41%, and a mean Average Precision (mAP) of 96.26%. The prediction model yielded an MAE of 18.11 and an  score of 0.53, indicating moderate predictive performance. Unit and usability testing demonstrated effectiveness and ease of use, highlighting its potential to improve the general quality of life.

Rsha Mirza, Dareen Alsulami, Amal Aljadani et al. · 0 citations