Evaluating Structured Pruning and Quantization Strategies for Edge-Efficient Traffic Object Detection
Abstract
Traffic object detection for deployment is constrained by accuracy, latency, model size, and computational cost. Under a unified dataset, deployment pipeline, and single-GPU hardware platform (RTX 4090), this paper evaluates the combined effects of structured pruning and low-precision inference. Using YOLOv5s as the primary model and YOLOv8n for transfer validation, it compares FP32, FP16, PTQ INT8, and QAT INT8 at pruning ratios of 0%, 30%, 50%, and 70% on BDD100K. The results show that structured pruning provides the dominant reductions in parameters and FLOPs; FP16 reduces latency with minimal accuracy loss; PTQ INT8 further reduces model size but yields larger accuracy drops; and QAT INT8 better preserves accuracy at similar INT8 cost. Within this dataset and hardware setting, a moderate pruning ratio around 50% offers a practical accuracy-efficiency operating point, while the lighter YOLOv8n is more sensitive to joint compression. The study provides a reproducible comparison baseline and deployment guidance for GPU-based traffic-scene detection models.