Skip to content
Open access

UAV acoustic recognition in complex environments based on attention mechanism fusion and variational mode decomposition

Aug 2026 · Engineering Research Express · Vol 8, pp. 165211 · 0 citations · 35 references
Physics

TL;DR

An improved deep learning model, ResNet18_Attention, is proposed based on the traditional ResNet18, which effectively enhances the feature representation capability of UAVs under low SNRs, and the pro posed model achieves consistent improvements in accuracy, precision, recall, and F1-score under noisy conditions.

Abstract

With the rapid development of unmanned aerial vehicle (UAV) technology, unauthorized small-drone flights are increasing, threatening civil aviation and public safety. Since radio frequency, electromagnetic and vision-based methods can be unreliable-especially in fog-acoustic UAV identification is a practical alternative. This study proposes a deep-learning method for UAV sound recognition and tests its real-world performance by mixing UAV audio with diverse background noises at multiple signal to noise ratio (SNR) levels to simulate realistic conditions.This study uses Mel-spectrograms as the extracted features for UAV audio recognition. By parallelly fusing squeeze-and-excitation networks and coordinate attention mechanisms, an improved deep learning model, ResNet18_Attention, is proposed based on the traditional ResNet18, which effectively enhances the feature representation capability. At the same time, the variational mode decomposition filtering algorithm is employed as a pre processing step to perform noise suppression before feature extraction, further improving the recognition and classification performance of UAVs under low SNRs. This study conducts a comparative evaluation between the proposed improved model, a model enhanced solely with a single attention mechanism, and four conventional baseline models. The pro posed model achieves consistent improvements in accuracy, precision, recall, and F1-score under noisy conditions. In addition, a filtering algorithm is employed to preprocess the UAV audio data.Experimental results show that, after filtering, the improved model achieves an average accuracy that is 15.1% higher than that of the ResNet18 model without filtering within the experimental SNR range, and the final average accuracy reaches 98.6% over this SNR range.

Read PDF

Similar papers

Open access Jul 2026

A BiLSTM and self-attention deep learning method for enhancing OFDM signal detection capability in unmanned swarms

Unmanned swarm systems—including unmanned aerial vehicles (UAVs), unmanned ground vehicles (UGVs), and autonomous surface vehicles—critically depend on highly reliable and secure communication networks for cooperative perception, collaborative decision-making, and coordinated control. Orthogonal Frequency Division Multiplexing (OFDM) technology is widely adopted in swarm communication due to its high spectral efficiency and strong resistance to multipath fading. However, traditional OFDM signal recognition methods heavily rely on manually extracted features (e.g., cyclic prefix, cyclic spectrum), and their recognition performance deteriorates sharply under low signal-to-noise ratio (SNR) and complex channel conditions—conditions commonly encountered in dynamic swarm environments with fast-varying topologies and electromagnetic interference. This paper proposes a deep learning-based OFDM signal detection method specifically designed for unmanned swarm communication scenarios. The proposed approach achieves higher recognition accuracy and stronger robustness under low SNR conditions compared to conventional techniques. By reducing dependence on prior knowledge and manual feature engineering, our method provides an effective solution for intelligent signal recognition in complex electromagnetic environments. Simulation results demonstrate that under low SNR conditions, across different subcarrier modulation schemes, the proposed model achieves lower channel estimation mean squared error (MSE) and superior bit error rate (BER) performance relative to traditional channel estimation techniques. These capabilities directly support the highly safe and reliable communication networks required for mission-critical swarm applications, including multi-UAV cooperative search, anti-terrorism operations, and disaster relief coordination.

Lijun Han, Feng Yi, Jing Wang · 0 citations
Open access Aug 2026

Range estimation of low-frequency underwater acoustic target based on deep learning architecture with data augmentation

Passive source localization is essential in underwater acoustics, yet both model‑based matched field processing (MFP) and conventional machine learning methods often suffer from limited accuracy and poor generalization. To address these challenges, this study presents a low-frequency underwater acoustic ranging approach that integrates data augmentation with the ResNet-UNet architecture. Using the real and imaginary components of the covariance matrix as inputs, the sample expansion is first performed by combining the deep convolutional generative adversarial network with several conventional augmentation techniques. Afterwards, a predictive model that fuses ResNet and U-Net is developed for target range estimation. The validity of the proposed method is examined through the SWellEX-96 sea trial data, where the performance is compared under two conditions, with and without the augmentation strategy, and also benchmarked against several reference methods, namely MFP, generalized regression neural network (GRNN), conventional convolutional neural network (CNN), ResNet, and the proposed ResNet-UNet. Experimental results indicate that the adopted augmentation can considerably enlarge the training sample set, which consequently enhances the ranging accuracy. The majority of MFP estimates fall beyond the acceptable error margin, while GRNN shows obvious weaknesses in generalization performance. Both the conventional CNN and ResNet are only capable of producing coarse range approximations. Nevertheless, when coupled with the proposed augmentation, the ResNet-UNet method effectively accomplishes range estimation and its performance markedly surpasses that of the other models. Moreover, it remains effective even under low low signal-to-noise ratio conditions.

Qi-Hai Yao, Zi-Jie Zhao, Jia-Xin Lu et al. · 0 citations
Preprint Aug 2026

Acoustic UAV Detection in Battlefield Scenarios: Handling Noise, Domain Shift, and Weak Labels

Passive acoustic sensing offers a critical, cost-efficient, and, crucially, passive alternative for detecting small unmanned aerial vehicles. However, the practical deployment of acoustic systems is discouraged by extreme environmental noise and sensor-induced domain shift caused by heterogeneous hardware. This paper addresses these challenges by introducing a robust framework optimized for real-world battlefield conditions. We propose the integration of Per-Channel Energy Normalization (PCEN) and attention-based pooling to enhance feature extraction under low signal-to-noise ratio scenarios. We further propose a domain-aware training strategy that leverages auxiliary classes and multi-microphone data to mitigate cross-domain performance degradation. Evaluated on a unique dataset of combat-zone recordings from the Ukrainian frontlines, our approach significantly outperforms existing baselines, increasing the F1 score from 55.4% to 78.6%. This paper was originally presented at the International Conference on Military Communication and Information Systems (ICMCIS), organized by the Information Systems Technology (IST) Scientific and Technical Committee, IST-224-RSY - the ICMCIS, held in Bath, United Kingdom, 12-13 May 2026.

Vadym Vilhurin, Volodymyr Sydorskyi, Andrii Shevtsov · 0 citations
Open access 2026

Radar-Audio UAV Classification Under Asymmetric Degradation Using Training-Inference Decoupled Gating

A controlled synthetic feature-level benchmark indicates that reliability-aware hard suppression can mitigate multimodal negative transfer under asymmetric degradation and should not be interpreted as evidence of robustness to waveform-level or real-world acoustic disturbances.

Chao Zhang, Xin Fang, Jingrui Zhang et al. · 0 citations
Open access Jul 2026

VA-DFN: An acoustic-vibration collaborative fusion network for bearings in strong noise environments

To address the limitation that a single sensor is insufficient for comprehensively extracting deep fault features in strong industrial noise environments, which constrains bearing diagnosis accuracy, this paper proposes an acoustic-vibration collaborative fusion network. First, an Adaptive Gated Residual Block (AGRB) is designed and combined with a Twin-Gated Residual Block (TGRB) architecture to effectively extract highly robust deep local acoustic and vibration features amidst strong background noise. Second, a Bidirectional Attention Sensing Module (BASM) is constructed to perform deep interaction and complementary calibration of heterogeneous acoustic-vibration features in the global semantic dimension, breaking through the limitations of traditional shallow concatenation of multimodal features. To verify the effectiveness of the proposed model, an experimental study was conducted on a 6205 deep groove ball bearing using a non-contact acoustic-vibration synchronous acquisition system with a 25 cm acoustic monitoring distance and a 5096 Hz sampling rate. The dataset contains nine diagnostic categories, including one healthy state and eight fault states.Experimental results indicate that this method can achieve deep dynamic alignment of heterogeneous data. The VA-DFN demonstrates exceptional noise-resistant robustness under varying signal-to-noise ratio (SNR) conditions from −6 dB to 2 dB, achieving a maximum diagnostic accuracy of 99.55%, which is significantly superior to existing single-modality and conventional deep learning baseline models.

Fanlong Zhu, Junyu Lai, Peiwen Lu et al. · 0 citations
Open access Aug 2026

Efficient and Interpretable Underwater Acoustic Target Recognition Using a Lightweight Heterogeneous Kernel Network

Underwater acoustic target recognition (UATR) is challenging due to the complex, multi-scale physical characteristics of marine targets and the strict computational limits of edge platforms like unmanned surface vehicles. To navigate the severe interference of underwater environments, existing methods increasingly rely on heavyweight architectures to achieve high recognition accuracy. However, the massive computational overhead of these models is fundamentally at odds with the restricted power and processing capabilities of practical deployment platforms. To resolve this conflict between performance and deployability, we propose LHK-Net, a lightweight Heterogeneous Kernel Network. By integrating a Heterogeneous Kernel Pyramid with Residual Depthwise Separable Convolutions, LHK-Net dynamically captures multi-scale acoustic features, from macroscopic steady-state harmonics to localized transient impulses, while compressing the model size to merely 0.82 M parameters. Additionally, a dual-domain Time–Frequency Attention module and an Adaptive SK-Fusion mechanism are incorporated for robust noise suppression. Experiments on the DeepShip dataset demonstrate that LHK-Net achieves state-of-the-art accuracy, outperforming heavyweight models at real-time speeds. Extensive visual analyses further validate that the network possesses strong physical interpretability, effectively aligning its internal feature representations with the intrinsic acoustic properties of the targets.

Yilling Sun, Menghao Fan, Haonan Wei et al. · 0 citations