Research on an Enhanced Zipformer With Multi-Feature Attention for Aviation Speech Recognition
Abstract
Automatic Speech Recognition (ASR) for Air Traffic Control (ATC) is challenging due to factors such as fast speech rates, significant background noise, domain-specific terminology, and limited annotated data. These factors complicate the development of models that are both highly reliable and low-latency. Although the Zipformer architecture has achieved promising performance in end-to-end ASR, its ability to effectively capture multi-scale contextual information and jointly model temporal-frequency dependencies remains limited in complex ATC communication scenarios. This paper introduces a efficientt end-to-end (E2E) Multi-Feature Attention ASR framework to tackle these challenges. Built upon an Enhanced Zipformer, the model integrates a Multi-Scale Parallel Convolution (MSPC) module and a Multi-Attention Module (MAM) for joint time-frequency modeling, while maintaining low-latency decoding efficiency. Evaluations on the self-constructed aviation speech dataset show that the proposed framework reduces the Character Error Rate (CER),using modified beam search, from 4.63% to 4.51% on the development set and from 4.87% to 4.73% on the test set, corresponding to relative reductions of 2.6% and 2.9%. The results highlight the practical potential of the proposed method for practical ATC speech recognition scenarios and provide a foundation for future work in multilingual adaptation and model quantization.