A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition
Automatic speech recognition is typically trained assuming that the reference transcript is the only valid labeling of an utterance, yet even nominally verbatim transcripts contain localized differences in pronunciation, spelling, or lexical realization that the acoustics do not uniquely determine. Omni-temporal Classi...