Skip to content

Quantifying Consonant Contributions to Word Intelligibility via Acoustic Masking

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

This paper presents a scalable method that measures consonant contribution using acoustic masking, and relates MMR to two linguistic factors previously reported to correlate with consonant contribution, namely phoneme frequency and functional load.

Abstract

Consonants contribute unequally to whether a word is understood. Given the limited time available for therapy, ranking consonants by contribution to intelligibility helps prioritize intervention targets in motor speech disorders. However, measuring this contribution relies on perceptual studies that are difficult to scale. This paper presents a scalable method that measures consonant contribution using acoustic masking. We silence one consonant at a time in an isolated word and test whether an automatic speech recognition (ASR) model still recognizes the word. We define a consonant's contribution score as the proportion of its masked instances for which the word becomes misrecognized, which we refer to as the mask-induced misrecognition rate (MMR). We relate MMR to two linguistic factors previously reported to correlate with consonant contribution, namely phoneme frequency and functional load. We apply this analysis across four languages, English, Spanish, German, and Czech, using three ASR architectures, MMS (encoder-only), Whisper (encoder-decoder), and Qwen3-ASR (LLM-based). Using partial Spearman correlations, we find that phoneme frequency correlates negatively with MMR while functional load correlates positively. In other words, more frequent consonants are less disruptive when masked, whereas consonants carrying more lexical contrast are more disruptive. Further cross-language analysis shows that consonant rankings agree only partially across languages, indicating that consonant contribution is language-dependent.

View source

Similar papers

Open access Sep 2026

Visual speech enhances phoneme separability in human superior temporal gyrus

Visual speech, such as lipreading, facilitates spoken word recognition, but the neural mechanisms underlying audiovisual speech perception remain poorly understood. Visual cues may disambiguate fine-grained articulatory features during early perceptual stages or instead integrate with speech at more categorical, phonem...

Yi-Ke Li, Iain DeWitt, Jonathan R. Brennan et al. · 0 citations
Preprint Sep 2026

Automated Assessment of L2 Speech Rhythm Using Low-Frequency Amplitude Modulations

Automated Speaking Assessment of non-native speech must effectively evaluate prosody, including speech rhythm, to align with human perception. However, commonly employed rhythm metrics rely on segmental duration, requiring an additional alignment step, which is error-prone in non-native speech containing disfluencies a...

João Lima, Lucas H. Ueda, P. Costa · 0 citations
Open access Sep 2026

Toolkit for acoustic–phonetic analysis of naturalistic speech data

This work demonstrates TAPA on the 2016 U.S. presidential debate, and suggests that TAPA can be used to increase access to naturalistic speech data and speed up the processing timeline with experts' supervision.

Ethan Kutlu, Emerson Peters, Ciara Tapanes et al. · 0 citations
Open access Aug 2026

The Effect of Plosive Content on the Loudness Perception of Vowel-Consonant-Vowel Syllables in Listeners with Sensorineural Hearing Loss

Objective This study investigated whether plosive consonants carry a perceptual loudness weighting that significantly exceeds that of non-plosive consonants when judged by hearing-impaired listeners. Design A prospective loudness matching experiment utilizing the method of adjustment. Study Sample 19 consenting native...

Thomas Davis, S. Bleeck · 0 citations

The role of segmental and tonal information in Thai spoken word recognition

An issue that continues to be debated in psycholinguistic research concerns the role of phonological information in spoken word processing, particularly segmental information (e.g., consonants and vowels) in comparison with suprasegmental information (e.g., lexical tone). While segmental information can be critical for...

Tanradee Mongkolpiyathana · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.