Toward Extremely Low-Bit and Multi-Precision Conformer and Speech Foundation Model Quantization
Abstract
Model quantization facilitates Automatic Speech Recognition (ASR) deployment on resource-constrained devices, yet existing methods suffer from severe accuracy loss below 4 bits, redundant storage for multiple precisions, and limited applicability across training paradigms. We address these challenges with two methods that share a hierarchical multi-precision design. When full retraining is feasible, Quantization-Aware Co-Training (QACT) jointly optimizes weight-shared 2-bit and 1-bit sub-networks. Its 2-bit Conformer achieves 12.86% WER on Switchboard, statistically equivalent to the FP32 baseline by MAPSSWE at $\alpha {=}0.05$, while supporting both precisions in one model. For foundation models where retraining is prohibitive, Codebook-Shared Vector Quantization (CSVQ) hierarchically optimizes shared codebooks at the post-training stage. CSVQ substantially improves over scalar Post-Training Quantization on HuBERT and Whisper and supports 2/3/4-bit inference without additional multi-precision storage. Evaluations across supervised Conformer, self-supervised Wav2Vec2 and HuBERT, and weakly-supervised Whisper show that QACT provides a high-accuracy solution when retraining is affordable, whereas CSVQ offers efficient compression under strict computational budgets.1