Acuity-DETR: Enhancing Perceptual Acuity for Small Object Detection in Driving Scenes
Abstract
Small objects in complex driving scenes often contain weak texture information and are highly sensitive to localization errors, which poses significant challenges to feature representation and bounding box regression. To address these issues, this paper proposes Acuity-DETR, a high-acuity visual perception model for small object detection in driving scenarios. First, a Dual-Domain Interaction Attention (DDIA) is designed to enhance the responses of potential target regions in the spatial domain and complement fine-grained small-object features with frequency-domain information. Second, a Multi-Scale Context Enhancement Module (MSCEM) is introduced to fuse local details and neighborhood contextual information through convolutional operations with different receptive fields, thereby alleviating the insufficient semantic representation of small objects. Finally, a Scale-Adaptive Cross-Layer Iterative Refinement (SACIR) is proposed, which uses the predicted box from the previous decoder layer as the spatial prior for the current layer and dynamically adjusts the bounding box refinement magnitude according to object scale, thereby improving the localization stability of small objects. Extensive experiments are conducted on the SODA-D, BDD100K, and Cityscapes datasets. On SODA-D, Acuity-DETR achieves 29.6% AP and 61.2% AP50 while maintaining low computational complexity. It further obtains 30.4% AP on BDD100K and 31.9% AP on Cityscapes, outperforming representative detection methods across different driving-scene distributions. Ablation studies and qualitative results demonstrate the effectiveness of Acuity-DETR in enhancing small-object features and optimizing localization.