This work evaluates two context biasing methods based on Whisper against three speech LLMs across read and non-read speech, reporting biased, unbiased, and overall word error rate (WER).
Abstract
Recognizing new and rare words - named entities, acronyms, domain specific special words, and other items scarce in training data - remains a key challenge for automatic speech recognition (ASR). We compare two strategies for this: context biasing methods, where an ASR model is extended such that during inference a word list can be supplied, and speech large language models (LLMs) prompted with context directly. We evaluate two context biasing methods based on Whisper against three speech LLMs across read and non-read speech, reporting biased, unbiased, and overall word error rate (WER). The context biasing methods cut biased WER by up to 88% relative while leaving other words largely unaffected. Speech LLMs excel on read speech but generalize less well to non-read speech, and prove sensitive to distractor count and prompt word order. We characterize the resulting trade-offs to guide method selection.
This work proposes Agentic-GER, an LLM-based agent for terminology correction in long-form speech, which uses global context from the full transcript to identify suspicious terms and resolve ambiguous hypotheses.
Yan-Qiao Zhu, Wu-Peng Wang, Zhifu Gao et al.· 0 citations
Recognition of domain-specific and low-frequency terms remains challenging for automatic speech recognition (ASR). Although contextual biasing can improve their recognition, directly providing a large terminology dictionary introduces many irrelevant biasing terms. Retrieval-based contextual biasing addresses this issu...
Natsuo Yamashita, Ai Nemoto, Ryosuke Koichi et al.· 0 citations
This work demonstrates that cross-lingual word-to-speech mappings can be learned directly from visual grounding without transcriptions or explicit model training.
Gabriel Pirlogeanu, Dan Oneata, H. Cucu et al.· 0 citations
It is shown that this simple approach to disentanglement can enhance higher-order linguistic information in word discovery tasks and learn representations which encode words with reasonable fidelity independently of local phonetic content.
A linguistically curated bilingual dataset comprising 13,846 utterances distributed across synthetic and literature-derived subsets, with OOV rates reaching up to 100%, and a diagnostic evaluation framework that partitions recognition performance into Biased Word Error Rate (B-WER), which targets domain-specific jargon...
Emily Haubert Klering, E. Cortes, Т. И. Черненко et al.· International Conference on...· 0 citations
It is demonstrated that the CTC-weighted language model log likelihood ratio can be interpreted as the KLD between the unknown correct ASR and the unsupervised learned ASR, and that, using a Pinsker bound, the square root of KLD can be interpreted as an upper bound on the total variation distance between the true and e...
Meng-Qi Wang, Mark A. Hasegawa-Johnson, Hao-Long Zheng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.