The central finding is that aggregate WER hides code switching behavior, and the best system by WER (an ASR model) is statistically indistinguishable from a leading audio LM on WER, yet the audio LM is significantly better on every switch localized metric.
Abstract
Automatic speech recognition (ASR) systems and audio language models (audio LMs) now report low error rates on monolingual benchmarks, but their behavior on code switched speech in low resource, diacritic rich languages remains poorly characterized. We present a switch aware evaluation of eleven modern systems (six ASR models and five audio LMs) on English Yoruba code-switched speech, using a deterministic 2000 utterance evaluation set and a shared scoring pipeline. Beyond word error rate (WER), we report switch localized diagnostics: a switch entry token error rate (SETER), windowed switch point error rates, language specific error rates, and a diacritic insensitive WER. Our central finding is that aggregate WER hides code switching behavior. The best system by WER (an ASR model) is statistically indistinguishable from a leading audio LM on WER, yet the audio LM is significantly better on every switch localized metric. Across faithful systems, Yoruba token recognition collapses (error 0.97 for almost all systems) while English tokens are recognized far better, and errors concentrate sharply at switches into Yoruba. Several generative audio LMs fail as exact transcribers, producing translation, verbosity, and prompt leakage that are strongly prompt dependent. We release manifests, metric implementations, and evaluation scripts to support reproducible, switch aware benchmarking for African code switched speech.
Forced alignment aligns speech audio with a text transcript to generate word and phone timestamps. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices once and releases the...
Wei Chu, Yuan-Zhe Dong, Ke Tan et al.· 0 citations
For automated speech recognition (ASR) systems, code-switched speech in which speakers alternate between two or more languages in a single utterance presents substantial difficulties, especially when it comes to Tamil-English languages. By creating a robust code-switched corpus and a parameter-efficient ASR system spec...
S. S., B. B· Journal on Audio, Speech, an...· 0 citations
This work investigates ESPnet-SpeechLM as a token-based backbone for generating SD hypotheses, formulating SD as autoregressive generation of structured tokens conditioned on acoustic input and shows that outputs generated by SpeechLMs encode useful temporal SD structure, but full-meeting SD remains limited by recordin...
This work evaluates two context biasing methods based on Whisper against three speech LLMs across read and non-read speech, reporting biased, unbiased, and overall word error rate (WER).
PURPOSE
Automated Speech Recognition (ASR) is used in assistive technologies such as caption phone and video calls for people who are deaf or hard of hearing. Evidence indicates that ASR's poor caption accuracy with "accented" speech (not the "standard" United States Broadcasting Mid-Western accent) is a barrier to suc...
Andrea Urqueta Alfaro, Karina A. Roundtree, Mark S. Pfaff et al.· Disability and Rehabilitatio...· 0 citations
Across systems, WER has little rank agreement with CTEM (Spearman $\rho=-0.28$) or TSR (Spearman $\rho=-0.28$), and even the strongest system leaves nearly one-third of recordings with an unrecovered critical value.
Tyler Baumgartner, Brandon Tai, Lisa Kaelin-Martin et al.· 0 citations