Interactive Multilevel Fusion With Dynamic Bounding Boxes for Small Object Detection in Clutter
Abstract
Detecting small targets in remote sensing imagery remains a critical challenge due to the inherent limitations of low-resolution targets and the frequent clustering of small targets. To address these challenges, we propose an integrated set of improvements to the You Only Look Once (YOLO)v8 framework, leveraging cross-connection mechanisms to construct a robust small target detector. Specifically, we first introduce a Multiscale Feature Enhancement Module (MFEM) that utilizes multiscale depth-separable convolutions combined with a cross-connection strategy. This design effectively enriches feature representations, significantly improving feature extraction for small targets while maintaining detection performance for medium to large targets. Second, we develop a multilayer feature fusion neck (MFFN) architecture that strategically integrates low-level spatial features into high-level semantic layers. This approach significantly improves detection performance in clustered small target scenarios while optimizing the overall feature fusion process. Third, to overcome the problem of limited adaptability of specialized small target losses in complex environments, we propose a dynamic normalized Wasserstein distance (DNWD) loss by combining the advantages of dynamically adjusted normalized Wasserstein distance (NWD) and complete intersection-over-union (CIoU). Extensive experiments demonstrate that our integrated approach consistently outperforms YOLOv8 baselines across multiple model scales. The proposed framework achieves mAP50 gains of 1.6% (YOLOv8n) and 1.8% (YOLOv8s) on the detection in optical remote sensing image (DIOR) dataset, as well as 2.1% (YOLOv8n) and 2.9% (YOLOv8s) on the VisDrone2019 dataset.