Toward Robust RF Drone Detection at Energy Facilities: A Synthetic Industrial-EMI Benchmark and Noise-Augmented Training Method
Abstract
Radio-frequency (RF) classifiers for drone detection at critical energy facilities reach 99%+ accuracy in laboratory conditions, but the electromagnetic interference (EMI) of substations, refineries, and nuclear plants is absent from public benchmarks. We show that exposing the classifier to a parameterized synthetic industrial-EMI distribution at training time recovers most of the lost accuracy at a small cost in clean-condition performance. We assemble a five-profile synthetic EMI model (AWGN baseline, Middleton Class-A impulsive, narrowband SCADA/PLC, partial discharge, and co-channel WiFi) with parameters drawn from the power-systems measurement literature, and benchmark eleven classical and deep learning detectors on the public DroneRF dataset across seven SNR levels and three classification tasks. All experiments use recording-level grouped data partitions—every source recording is assigned entirely to one of the training, validation, or test sets before window selection—and each augmented model shares identical hyperparameters with its non-augmented baseline, so the comparisons isolate augmentation from both data leakage and model capacity. Under this protocol, noise-augmented training lifts binary-detection accuracy at SNR = 0 dB by up to 27.5 percentage points under SCADA interference (matched 1D-CNN-V2 backbone) and 23.9 percentage points under partial discharge for the capacity-matched random-forest pipeline (Holm–Bonferroni-corrected $p \le 0.011$ ), with average gains of 6–10 percentage points over the full −10 to + 20 dB sweep; the gains are stable across five independent grouped splits. Augmentation adds no inference-time cost. A leave-one-noise-out analysis shows the gains transfer fully to unseen impulsive interference and partially to broadband noise, but not to structurally distinct narrowband SCADA, and an EMI-only operating-characteristic study shows augmentation lowers both the per-window false-positive and miss probabilities, with zero event-level errors under temporal voting on held-out recordings. Synthetic-only EMI remains a deliberate limitation, with field validation the immediate next step; we release code and noise generators to support it.