A controlled synthetic feature-level benchmark indicates that reliability-aware hard suppression can mitigate multimodal negative transfer under asymmetric degradation and should not be interpreted as evidence of robustness to waveform-level or real-world acoustic disturbances.
Abstract
Multimodal radar–audio sensing combines the spatial and reflectivity cues provided by millimeter-wave radar with the class-discriminative spectral patterns of acoustic signals for unmanned aerial vehicle classification. Under asymmetric degradation, however, a corrupted acoustic branch may become detrimental because continuously weighted acoustic features can remain in the fused representation. To investigate this failure mode, a controlled benchmark is constructed by applying Gaussian perturbations to stored Log-Mel representations while keeping the paired radar inputs unchanged. RAD-Gating is proposed as a reliability-aware fusion framework that decouples classifier-backbone training, reliability-gate calibration, and threshold-based inference. After the classifier backbone has been trained, its parameters are frozen and a lightweight reliability estimator is calibrated using polarized clean and degraded samples. During inference, the predicted acoustic reliability score is converted into a binary gate; unreliable acoustic features are zero-masked, and activation-scale compensation is applied to the retained radar representation. On the fixed 80/20 MMAUD split, the accuracy at −10 dB controlled feature-level degradation is increased from 29.23% with Softmax fusion to 56.15% with RAD-Gating, while the dedicated radar-only classifier achieves 61.54%. Fixed-checkpoint analyses are further conducted to examine reliability-score separation, threshold selection, feature-scale control, and soft versus hard inference. Within this controlled synthetic feature-level benchmark, the results indicate that reliability-aware hard suppression can mitigate multimodal negative transfer. These findings should not be interpreted as evidence of robustness to waveform-level or real-world acoustic disturbances.
Experimental results demonstrate that the proposed sequential pipeline preserves high specificity while reducing missed detections compared to radar-only processing and indicates that sequential processing can offer a viable alternative to parallel fusion for edge-based drone detection under SNR and computing constraints.
Faizal Mohd Amin Sharifuldin, N. E. Abdul Rashid, Mohd Adli Md Ali et al.· Sensing and Imaging· 0 citations
Passive acoustic sensing offers a critical, cost-efficient, and, crucially, passive alternative for detecting small unmanned aerial vehicles. However, the practical deployment of acoustic systems is discouraged by extreme environmental noise and sensor-induced domain shift caused by heterogeneous hardware. This paper addresses these challenges by introducing a robust framework optimized for real-world battlefield conditions. We propose the integration of Per-Channel Energy Normalization (PCEN) and attention-based pooling to enhance feature extraction under low signal-to-noise ratio scenarios. We further propose a domain-aware training strategy that leverages auxiliary classes and multi-microphone data to mitigate cross-domain performance degradation. Evaluated on a unique dataset of combat-zone recordings from the Ukrainian frontlines, our approach significantly outperforms existing baselines, increasing the F1 score from 55.4% to 78.6%. This paper was originally presented at the International Conference on Military Communication and Information Systems (ICMCIS), organized by the Information Systems Technology (IST) Scientific and Technical Committee, IST-224-RSY - the ICMCIS, held in Bath, United Kingdom, 12-13 May 2026.
Passive acoustic target recognition is often constrained by the complex interplay of variable underwater propagation channels and nonstationary target states. While data-driven deep learning models offer exceptional flexibility in feature learning, they are frequently susceptible to overfitting environmental noise and site-specific cues, which undermines their generalization in fluctuating marine conditions. Conversely, methods grounded in physical attributes exhibit superior intrinsic stability across diverse environments but typically lack the comprehensive signal perception and discriminative richness required for sophisticated classification. To bridge this gap, we propose rhythm-aware adaptive spectro-temporal enhancement (RASTE), prioritizing physical interpretability and robustness. Unlike traditional detection of envelope modulation on noise methods that require manual bandpass filter selection and assume signal stationarity, RASTE adaptively extracts rhythmic signatures and maintains efficacy, even under non-stationary conditions, such as pulsed interference. These features are applied as a soft mask to spectrograms to integrate physical priors while preserving the integrity of discriminative features. Experiments on open-source datasets indicate that RASTE serves as a robust and interpretable alternative to baselines. By navigating the performance-interpretability trade-off, RASTE achieves competitive results, particularly in scenarios characterized by pronounced rhythmic structures.
Zhengkun Liu, Jiawei Ren, Ji Xu et al.· Journal of the Acoustical So...· 0 citations
An improved deep learning model, ResNet18_Attention, is proposed based on the traditional ResNet18, which effectively enhances the feature representation capability of UAVs under low SNRs, and the pro posed model achieves consistent improvements in accuracy, precision, recall, and F1-score under noisy conditions.
Jiajun Huang, Kuangang Fan, Zhiyu Zeng et al.· Engineering Research Express· 0 citations
Indoor human activity recognition using radar is a promising approach for privacy-preserving monitoring in assisted living and safety applications. However, radar-based recognition remains sensitive to viewpoint geometry because the same activity can produce different motion patterns when observed from different sensor positions. This work investigates whether compact per-radar classifiers and simple decision-level fusion are sufficient for reliable indoor human activity recognition using three frequency-modulated continuous-wave radars installed at different elevations. Point cloud detection files exported from the radar are converted into Doppler–Range, Doppler–Time, and Range–Time domain maps. A lightweight separable convolutional neural network is trained independently for each radar viewpoint, and session-level probability vectors are combined using mean fusion, entropy-weighted fusion, confidence-weighted fusion, and Dempster–Shafer fusion. Fusion is evaluated with session-level five-fold cross-validation, so every fused prediction comes from models that never saw that session during training. The per-radar probabilities are also temperature calibrated to check that the comparison between fusion rules is fair. The results show that viewpoint geometry strongly affects single-radar performance. Under leave-one-subject-out (LOSO) evaluation, the two frontal radars achieve macro F1 scores of 0.771 and 0.728, while the downward-angled radar reaches 0.480. Domain-map ablation shows that the Range–Time channel is the least informative map on every viewpoint. Differences among the stronger Doppler-based combinations stay within run-to-run training variation. Decision-level fusion gives the strongest session-level performance. The frontal pair reaches a macro F1 score of 0.952 with simple mean fusion and 0.964 with Dempster–Shafer fusion across 173 matched sessions. McNemar tests find no significant difference between mean fusion and the more complex rules. Fusing a strong radar with a much weaker viewpoint can even perform worse than the stronger radar alone.
Abdillah Nur Isnaini, F. Y. Suratman, Khilda Afifah et al.· IEEE Access· 0 citations
Continuous traffic monitoring is critical for accurate structural load assessment and fatigue life estimation of highway bridges. However, conventional vision-based methods suffer from limitations such as line-of-sight restrictions, susceptibility to adverse weather and lighting conditions, and limited spatial coverage. Distributed acoustic sensing (DAS) offers a robust alternative by repurposing existing telecommunications dark fiber into dense, kilometer-scale sensor arrays. Nevertheless, interpreting the complex DAS signals generated by vehicular traffic remains challenging due to overlapping dynamic signatures and the scarcity of ground-truth data for model training. To overcome this, we present a cross-modal (vision-to-optic) supervision framework that transforms existing fiber infrastructure into a traffic monitoring system. We deployed a synchronized camera–DAS testbed along a roadway segment served by dark fiber. Video data is processed using modern computer vision foundation models SAM3 to automatically extract vehicle trajectories and classifications. These camera-derived labels supervise a deep sequence learning model trained on the corresponding DAS strain data. Once trained, the fiber-optic system independently achieves accurate vehicle detection, classification, localization, and speed estimation. Finally, we demonstrate how these continuous, DAS-derived traffic metrics can be directly translated into dynamic load profiles, providing a scalable, continuous monitoring solution for bridge fatigue and structural health assessment.
Cong Chen, Shenghan Zhang· e-Journal of Nondestructive...· 0 citations