Skip to content

Author

Mohamed Oubenal

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

An Integrated Semi-Supervised Framework for Low-Resource Language ASR: Acoustic Enhancement, LoRA Fine-Tuning, and Hybrid Post-Processing for Amazigh Variants

Low-resource languages continue to face significant challenges in automatic speech recognition (ASR), especially those with limited annotated corpora and significant variant variation. The Amazigh language family, which is spoken throughout North Africa, has very little digital infrastructure and is severely underfunded. In this paper, we present a multi- variant Amazigh ASR system that integrates linguistically-informed post-processing correction and audio enhancement preprocessing with refined multilingual self-supervised models is presented. We address three key issues: (1) acoustic degradation in field recordings; (2) substantial inter-variant variation; and (3) extreme data scarcity. Our enhanced pipeline adds a semi-supervised learning framework that includes: (1) data augmentation through synthetic speech generation and aggressive spectral perturbation, which increases the training corpus to 15 hours; (2) self-training on unlabeled data using pseudo-labeling; (3) VoiceFixer-based audio restoration; and (4) hybrid n-gram/Levenshtein post-correction. WER of 18.7% ± 2.3% (95% CI), CER of 9.4% ± 1.6%, and PER of 12.1% ± 1.9% show statistically significant improvements, with a relative WER reduction of 34.2% compared to baseline Whisper (p < 0.001). Component contributions are quantified by ablation studies: audio preprocessing results in a relative improvement of 8.3%, while post-processing adds 12.7% reduction. In addition to offering a broadly applicable framework for the preservation of low-resource languages, this work sets new state-of-the-art for Amazigh ASR.

Youness Chaabi, Mohamed Oubenal · 0 citations