Skip to content

Author

Yuanrong Shen

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Aug 2026

Using accent variability to probe the performance of the Whisper automatic speech recognition system

Automatic speech recognition (ASR) systems often achieve high accuracy for native speech, yet remain less reliable for non-native (L2) accented speech. This gap raises a question about why ASR performs so well on L1 speech. When acoustic cues diverge from expectation, does ASR accommodate L2 speech acoustics, or does it rely on semantic predictability to infer likely words? To separate acoustic sensitivity from semantic inference, this study uses Whisper to probe how model size, semantic context, and talker-level variability influence transcription accuracy under accent-related variability. Five Whisper models (tiny, base, small, medium, and large-v3) were used to transcribe 200 read sentences, half high-predictability and half low-predictability, produced by 24 L1 Mandarin speakers of English and 24 L1 American-accented English speakers. We characterize each talker’s vocabulary knowledge, accent exposure, and perceived accentedness. Mixed-effects analyses of transcription accuracy will test three predictions. First, increasing model size will predict higher accuracy for both L1 and L2 speech, with an outsized effect for L2 speakers. Second, high-predictability sentence context will predict higher target-word accuracy, with a larger benefit for L2 speech. Third, continuous measures of lexical proficiency, prior accent exposure, and perceived accentedness will explain transcription accuracy beyond a categorical L1–L2 distinction.

Yuanrong Shen, Oishani Bandopadhyay, Sarah C. Creel · 0 citations