Skip to content
Open access

Toward Extremely Low-Bit and Multi-Precision Conformer and Speech Foundation Model Quantization

2026 · IEEE Transactions on Audio, Speech, and Language Processing · Vol 34, pp. 4331-4347 · 0 citations · 65 references

Abstract

Model quantization facilitates Automatic Speech Recognition (ASR) deployment on resource-constrained devices, yet existing methods suffer from severe accuracy loss below 4 bits, redundant storage for multiple precisions, and limited applicability across training paradigms. We address these challenges with two methods that share a hierarchical multi-precision design. When full retraining is feasible, Quantization-Aware Co-Training (QACT) jointly optimizes weight-shared 2-bit and 1-bit sub-networks. Its 2-bit Conformer achieves 12.86% WER on Switchboard, statistically equivalent to the FP32 baseline by MAPSSWE at $\alpha {=}0.05$, while supporting both precisions in one model. For foundation models where retraining is prohibitive, Codebook-Shared Vector Quantization (CSVQ) hierarchically optimizes shared codebooks at the post-training stage. CSVQ substantially improves over scalar Post-Training Quantization on HuBERT and Whisper and supports 2/3/4-bit inference without additional multi-precision storage. Evaluations across supervised Conformer, self-supervised Wav2Vec2 and HuBERT, and weakly-supervised Whisper show that QACT provides a high-accuracy solution when retraining is affordable, whereas CSVQ offers efficient compression under strict computational budgets.1

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.