Bridging the Scale Gap: A Multi-Scale Feature Enhancement Framework for UAV Aerial Image Object Detection
Abstract
Highlights What are the main findings? A UAV-specific RT-DETR framework, termed MSF-DETR, is developed to improve feature representation and query optimization for aerial object detection. Experiments on the DIOR and DOTA datasets show that MSF-DETR consistently outperforms RT-DETR and representative baseline detectors, achieving mAP50 scores of 86.3% and 77.2%, respectively. What are the implications of the main findings? The results demonstrate that coordinated cross-scale feature enhancement, channel discrimination, spatial-detail preservation, and query supervision are effective for challenging UAV scenes involving small and densely distributed targets. The proposed design provides an effective balance between detection performance and computational efficiency, indicating its potential for practical UAV remote-sensing applications. Abstract Unmanned aerial vehicle (UAV) imagery is a core data source for remote sensing interpretation, intelligent transportation, urban monitoring, and disaster assessment, yet its large scale variation, dense object distributions, and complex backgrounds continue to challenge automated detection systems. Transformer-based detectors offer strong global modeling capacity, but existing implementations still suffer from insufficient multi-scale feature interaction, weak discriminative representation, and loss of fine-grained spatial detail, which together limit performance on small and densely arranged targets. This paper proposes MSF-DETR, a multi-scale feature enhancement framework built on RT-DETR that integrates four coordinated components: an Enhanced Feature Connection (EFC) module for adaptive cross-scale interaction, a Feature Channel Attention (FCA) module for frequency-domain discriminative enhancement, a Reinforced Attention Feedback Module (RAFM) for spatial-detail preservation within the Transformer encoder, and a Unified Query Supervision Loss (UQSL) for stable dense-scene supervision. On the DIOR benchmark, MSF-DETR achieves 86.3% mAP50 and 64.4% mAP50–95, improving on the RT-DETR baseline by 2.8 and 2.6 percentage points, respectively; on DOTA, it reaches 77.2% mAP50 and 48.8% mAP50–95, improvements of 4.7 and 4.6 points. These results demonstrate that jointly coordinating multi-scale fusion, channel discrimination, spatial-detail retention, and query-level supervision, rather than stacking independent modules, yields measurable robustness gains for small and densely distributed objects in UAV aerial imagery.