An Integrated Semi-Supervised Framework for Low-Resource Language ASR: Acoustic Enhancement, LoRA Fine-Tuning, and Hybrid Post-Processing for Amazigh Variants
Abstract
Low-resource languages continue to face significant challenges in automatic speech recognition (ASR), especially those with limited annotated corpora and significant variant variation. The Amazigh language family, which is spoken throughout North Africa, has very little digital infrastructure and is severely underfunded. In this paper, we present a multi- variant Amazigh ASR system that integrates linguistically-informed post-processing correction and audio enhancement preprocessing with refined multilingual self-supervised models is presented. We address three key issues: (1) acoustic degradation in field recordings; (2) substantial inter-variant variation; and (3) extreme data scarcity. Our enhanced pipeline adds a semi-supervised learning framework that includes: (1) data augmentation through synthetic speech generation and aggressive spectral perturbation, which increases the training corpus to 15 hours; (2) self-training on unlabeled data using pseudo-labeling; (3) VoiceFixer-based audio restoration; and (4) hybrid n-gram/Levenshtein post-correction. WER of 18.7% ± 2.3% (95% CI), CER of 9.4% ± 1.6%, and PER of 12.1% ± 1.9% show statistically significant improvements, with a relative WER reduction of 34.2% compared to baseline Whisper (p < 0.001). Component contributions are quantified by ablation studies: audio preprocessing results in a relative improvement of 8.3%, while post-processing adds 12.7% reduction. In addition to offering a broadly applicable framework for the preservation of low-resource languages, this work sets new state-of-the-art for Amazigh ASR.