Bridging the Synthetic-to-Real Gap in Impulsive Sound Detection Using Audio Transfer Learning
Abstract
A CNN achieving 98.4% F1 on synthetic benchmark spectrogram data collapses to 20.0% F1 on real-world audio, revealing a severe synthetic-to-real domain gap in impulsive sound detection. This paper provides one of the first quantitative studies of this phenomenon and demonstrates that representation choice dominates classifier architecture when moving from curated datasets to real acoustic environments. We present a two-stage detection pipeline combining an energy-based prefilter (Stage 1) with a deep-learning classifier (Stage 2). Three Stage 2 architectures are evaluated: (a) an EfficientNetB0 CNN trained on synthetic spectrogram images using FFT, Log-Mel, and MFCC features; (b) a lightweight classification head trained on frozen YAMNet embeddings; and (c) a real-world fine-tuned CNN. Fine-tuning improves real-world F1 from 20.0% to 88.9%, while YAMNet achieves $98.07 \pm 0.24 {\%}$ F1 across three independent random seeds without synthetic pretraining, indicating a substantial advantage for audio-native representations. We state the scope of this evidence precisely: the 78.4 -point collapse is a genuine cross-corpus measurement, whereas the YAMNet figure is a held-out within-corpus result and should be read as an upper bound pending cross-corpus validation. The full system runs in real time on a standard laptop, supports a graphical dashboard, and is designed for deployment on edge devices.