A Dual-Scale Collaborative Vision Framework for UAV-Based Drowning Behavior Recognition
Abstract
To support the early identification of potential drowning-risk states and improve rescue response efficiency, this paper proposes a UAV-oriented dual-scale detection-pose cascade for frame-level drowning-risk recognition. The framework first performs high-recall preliminary detection on wide-field input images to identify potential drowning targets. The detected target regions are subsequently extracted and resized to construct localized inputs for the second-stage pose-based verification. This software-based target-region refinement simulates the localized high-resolution observation that could be provided by a telephoto camera in a future physical dual-camera UAV implementation. By focusing subsequent analysis on the localized target regions, the second-stage pose model can exploit finer-scale human structural information for drowning-risk state verification, thereby providing decision support for potential drowning detection. To address the challenges of small target scales and severe background interference in wide-field images, a lightweight YOLOv8n-based detection model is developed. An enhanced edge-feature-guided residual convolutional block attention module (EGRCBAM) is introduced, together with a recall-oriented FPIoU loss function designed for hard sample optimization, improving the recall of the drowning category by 17%. For localized target verification, an enhanced YOLOv8n-Pose model is constructed by incorporating a coordinate-aware pose head, spatial attention mechanism, and skeletal structure constraints to enhance human-region localization and pose-based drowning-versus-swimming recognition. The model improves Box mAP@0.5 from 0.768 to 0.816. Comparative experiments on the self-collected dataset demonstrate the effectiveness of the proposed detection and pose-based verification framework. The proposed framework provides a lightweight vision-based solution or UAV-oriented frame-level drowning-risk recognition and offers a potential algorithmic basis for future integration with physical dual-camera UAV platforms.