Skip to content

Learning New Words from Unlabeled Test Data in Automatic Speech Recognition

Sep 2026 · 0 citations · 31 references
Engineering Computer Science

TL;DR

It is demonstrated that the CTC-weighted language model log likelihood ratio can be interpreted as the KLD between the unknown correct ASR and the unsupervised learned ASR, and that, using a Pinsker bound, the square root of KLD can be interpreted as an upper bound on the total variation distance between the true and estimated spelling of the unknown word.

Abstract

New words are invented every day. A human listener can learn a new word by hearing it clearly once and inferring its usage from sentence context. This paper proposes granting ASR a similar ability to learn the contextual representations and spellings of new words from unlabeled test data at test time. A frozen CTC acoustic model provides spellings, a frozen language model provides contextual evidence for out-of-vocabulary (OOV) word detection, and an adaptation module expands the vocabulary by learning the lexical token representations with distributions over CTC-generated candidates. The spelling model of each token is optimized by minimizing a Kullback-Leibler divergence (KLD) objective. We demonstrate that the CTC-weighted language model log likelihood ratio can be interpreted as the KLD between the unknown correct ASR and the unsupervised learned ASR, and that, using a Pinsker bound, the square root of KLD can be interpreted as an upper bound on the total variation distance between the true and estimated spelling of the unknown word. Experiments show relative OOV character-error-rate reductions of up to 14.97% on LibriSpeech and 6.67% on dysarthric Speech Accessibility Project data for recurring OOV words, relative to the corresponding rescoring system.

View source

Similar papers

#natural language process... Preprint Sep 2026

A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition

Automatic speech recognition is typically trained assuming that the reference transcript is the only valid labeling of an utterance, yet even nominally verbatim transcripts contain localized differences in pronunciation, spelling, or lexical realization that the acoustics do not uniquely determine. Omni-temporal Classi...

Saurabh Kumar, Diptiman Mohanta, P. Ghosh · 0 citations
Preprint Aug 2026

Cached LLM Probability Retrieval for Speech Recognition

Large language models (LLMs) enhance automatic speech recognition (ASR) by providing linguistic priors; however, their direct rescoring is costly because it requires evaluating every N-best hypothesis. This paper introduces"cached LLM probability retrieval,"which involves querying a local teacher LLM offline to obtain...

Sheng Li, Takahiro Shinozaki, Tatsuya Kawahara · 0 citations
Preprint Sep 2026

Training-Free Contextual ASR via SpeechLLM-Based Error-Aware Selective Retrieval

Recognition of domain-specific and low-frequency terms remains challenging for automatic speech recognition (ASR). Although contextual biasing can improve their recognition, directly providing a large terminology dictionary introduces many irrelevant biasing terms. Retrieval-based contextual biasing addresses this issu...

Natsuo Yamashita, Ai Nemoto, Ryosuke Koichi et al. · 0 citations
2026

A Dataset for Evaluating ASR on Specialized Vocabulary

A linguistically curated bilingual dataset comprising 13,846 utterances distributed across synthetic and literature-derived subsets, with OOV rates reaching up to 100%, and a diagnostic evaluation framework that partitions recognition performance into Biased Word Error Rate (B-WER), which targets domain-specific jargon...

Emily Haubert Klering, E. Cortes, Т. И. Черненко et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.