MRAN-UNet: Physics-Informed Harmonic Frequency Attention for Multilingual Speech Enhancement
Abstract
Deep learning speech enhancement models are trained without grounding in acoustic physics, and evaluations remain confined almost exclusively to English. We address both gaps with MRAN-UNet, which embeds Harmonic Frequency Attention (HFA) - a parameter-free module derived from the source-filter model that aggregates spectral features at candidate $F_{0}$ positions and their harmonic overtones. On VoiceBank-DEMAND, MRAN-UNet achieves CSIG 4.77 (the highest among compared CNN/UNet/RNN baselines), STOI 0.927, and RTF 0.24 with only 3.1 M parameters. PESQ (2.42) trails the strongest convolutional baseline due to decoder spectral coloration, not the HFA mechanism - an effect confirmed by ablation. Complementing the architecture, we release Vaakdhara-DLSE-TE, the first paired enhancement corpus for Telugu (32,000 utterances). Zero-shot transfer improves Telugu STOI from 0.65 to 0.74; 20epoch fine-tuning reaches PESQ 1.91 and STOI 0.92 at 20 dB SNR, outperforming zero-shot DCCRN by 0.56 PESQ at 20 dB SNR.