A lightweight transformer-based YOLO for object detection in complex traffic scenarios
Abstract
Object detection is a fundamental perception task in intelligent transportation systems and autonomous driving. These systems rely on computer vision techniques to enable intelligent perception in complex traffic environments. However, object detection models still struggle in real-world driving scenarios, particularly when detecting occluded, small, and distant objects. To address these issues, this paper proposes YOLO-CasViT, a lightweight object detector built upon YOLOv8. Specifically, the traditional bottleneck structure in the C2f module is replaced with a Convolutional Additive Self-Attention block, forming a lightweight C2f_CasViT module in the backbone, driven by our data-informed module placement. This design enhances both local feature extraction and global contextual modeling. Experimental results on the KITTI dataset demonstrate that YOLO-CasViT improves precision, recall, mAP@0.5, and mAP@0.5:0.95 by 2.78%, 4.02%, 3.13%, and 0.49%, respectively, compared with the baseline. The results indicate that the proposed module improves detection performance in complex traffic scenarios and achieves competitive performance compared with representative detection models.