Skip to content
Open access

AMGAN: Adversarial Distribution Alignment of Acoustic Posteriors for Noise-Robust Automatic Speech Recognition

2026 · IEEE Access · Vol 14, pp. 123950-123964 · 0 citations · 35 references

Abstract

Automatic speech recognition (ASR) systems trained on clean speech often experience significant performance degradation when deployed in noisy acoustic environments. Existing approaches to improving robustness, such as speech enhancement and feature normalization, mainly operate at the signal or feature level and do not directly address the sensitivity of the acoustic model itself to noise. In this paper, we introduce AMGAN, a generative adversarial acoustic modeling framework designed to improve noise robustness at the acoustic model level. In the proposed approach, a clean-trained acoustic model is used as the generator to produce phoneme posterior distributions from noisy speech, while a discriminator encourages these outputs to align with oracle posteriors obtained from clean speech. In contrast to conventional GAN-based enhancement methods and teacher–student approaches that typically rely on point-wise supervision, AMGAN performs alignment at the distribution level in the posterior-probability space, allowing the model to learn representations that are more stable under noisy conditions. Experimental results on the TIMIT dataset show consistent improvements across MLP, RNN, and LSTM acoustic models, with the LSTM achieving absolute WER reductions of 3.4% and 2.7% under white and babble noise, respectively. On the larger LibriSpeech corpus, AMGAN outperforms both clean-only and multi-condition training baselines while maintaining competitive performance on clean speech. Furthermore, experiments with a pretrained Whisper encoder demonstrate that the proposed framework can be applied to modern end-to-end ASR systems, yielding an average absolute WER reduction of 2.3% under noisy conditions. Overall, the results suggest that adversarial alignment in posterior-probability space provides an effective and scalable way to improve ASR robustness without modifying the input representation, offering a practical alternative to enhancement-based methods.

Read PDF