UAV-LiteDet: A Lightweight Small Object Detection Network for Low-Altitude UAV Scenarios
Abstract
Vehicles and pedestrians in low-altitude UAV images usually present features such as small object scales, dense distribution, complex background textures and low contrast. Existing high-precision object detection models often have problems including a large number of parameters, high computational cost and difficulties in edge deployment. To address the above issues, this paper proposes a lightweight Anchor-Free small object detection network UAV-LiteDet for low-altitude UAV scenarios. The network uses depthwise separable convolution to build a lightweight backbone network, and designs a shallow cross-scale fusion module to fuse shallow detail features with stride 4 and deep semantic features with stride 8, so as to reduce the spatial information loss of small objects caused by continuous downsampling. At the detection end, an Anchor-Free grid prediction method is adopted to directly regress object confidence, category, center position and bounding box parameters, which avoids anchor box clustering and complex hyperparameter design. To reduce the sensitivity of small object center positions to single grid matching, this paper further introduces a Gaussian quality assignment strategy, which generates a smooth confidence supervision signal according to the distance between the grid and the object center, enabling grids near the object center to jointly participate in feature learning. Meanwhile, an area-adaptive weight is introduced into the bounding box regression loss to improve the optimization priority of small-scale objects during training. To lower the cost of data collection and manual annotation, this paper constructs the Synthetic-UAV-LowAlt synthetic low-altitude UAV small object dataset to simulate typical scenarios such as roads, buildings, ground textures, random illumination, noise, low contrast, as well as vehicles and pedestrians. Experimental results show that UAV-LiteDet has sound training convergence and high detection accuracy, and can stably identify small-scale vehicles and pedestrians in low-altitude UAV images under complex backgrounds. Compared with the baseline network without shallow cross-scale fusion, the proposed method shows obvious advantages in precision, recall, F1-score and object localization quality, while maintaining a small number of model parameters and fast inference speed. The detection result visualization and confusion matrix further indicate that the network has strong small object feature expression ability and class discrimination ability, and can well balance detection performance, computational complexity and real-time deployment requirements.