Skip to content
Conference

Adapting Whisper Models Using LoHA for Robust Recognition of Children's Speech

Jul 2026 · International Conference on Signal Processing and Communications · pp. 1-5 · 0 citations · 23 references

Abstract

Fine-tuning of large pre-trained models, such as Whisper, has gained prominence in the area of speech processing. Since full fine-tuning requires a large amount of data as well as high-ended computational resources, parameter efficient fine-tuning (PEFT) has been the preferred choice among researchers. During the past few years, several PEFT techniques have been developed and have been observed to be extremely effective. Motivated by the success of PEFT, in this paper, we have investigated and documented the efficacy of Low-Rank Hadamard Product Adaptation (LoHA) of Whisper models for children's automatic speech recognition (ASR) task especially in limited data scenario. We have also compared LoHA with a few other PEFT variants. Even though LoHA has been explored for signal processing and federated learning tasks, its impact on children's ASR has not yet been studied. Experimental results presented in this paper indicate that applying LoHA significantly influences ASR performance, achieving a word error rate of 3.0%.

View source