Generative Data Augmentation Method for Sonar Images Based on Diffusion Model
Abstract
Sonar object detection is constrained by scarce instance-level annotations, high acquisition costs, and long-tailed category distributions in real underwater environments. To address these limitations, this paper proposes a generative data augmentation framework based on the Stable Diffusion Model (SDM) for synthesizing sonar images together with target bounding boxes. The framework first fine-tunes SDM with an instance-level slice cropping strategy to strengthen the alignment between text prompts and local acoustic target structures. It then introduces a cross-modal sparse localization module (CMSL), which uses denoising features and text priors to infer 2D bounding boxes for generated samples. Synthetic long-tail samples are mixed with real URPC2022 data under a fixed-ratio saturation compensation strategy and evaluated through UTD-SCnet fine-tuning. The results show that instance-level cropping provides the best generation quality among the tested strategies (FID = 31.42, DR = 0.815, CCSCR = 0.075), and that a 60% synthetic-data injection ratio yields the best detection performance. These findings indicate that diffusion-based augmentation can provide a practical, semi-automated supplement for long-tail sonar detection, while excessive synthetic data may introduce domain-shift effects.