A linguistically structured approach to non-canonical phoneme recognition that decomposes phoneme prediction into articulatory feature dimensions such as manner, place, and voicing is introduced and suggests that articulatory feature supervision is a promising strategy for robust and interpretable phoneme recognition in non-canonical speech.
Abstract
Pathological and more broadly non-canonical speech present significant challenges for automatic phoneme recognition due to systematic deviations from canonical pronunciation and limited availability of labeled clinical speech data. Existing phoneme recognition systems are typically trained on canonical speech and treat phonemes as atomic categorical labels, limiting their ability to detect structured articulatory errors common in speech disorders and accents. In this work, we introduce a linguistically structured approach to non-canonical phoneme recognition that decomposes phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. We implement this formulation using a hierarchical multi-task learning architecture in which task-specific articulatory feature heads learn feature-level representations that are subsequently integrated through a cross-attention-based fusion module to produce phoneme predictions. To address the scarcity and noise of pathological speech labels, we combine this framework with semi-supervised learning via Momentum Pseudo-Labeling (MPL) and propose a cascaded training strategy that progressively introduces articulatory feature tasks while employing staged unfreezing of a pretrained speech encoder. Experiments on L2-ARCTIC, used as a proxy for pathological speech variation, show that the proposed approach achieves substantial improvements in phoneme recognition performance compared to strong baseline architectures, while yielding interpretable error patterns aligned with phonological feature structure. These results suggest that articulatory feature supervision is a promising strategy for robust and interpretable phoneme recognition in non-canonical speech, and motivate future validation on clinically diagnosed pathological speech datasets.
Preliminary analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation, indicating that phonological information from synchronized audio can be partially transferred to articulatory models.
Abner Hernandez, T. A. Vergara, Dai-Qi Liu et al.· 0 citations
The task of automatic speaker profiling based on speech signals becomes increasingly crucial in human-computer interaction, clinical voice assessment, and affective computing. Still, the tasks of speaker age group classification and emotion recognition are usually studied separately despite similar acoustic characteris...
R. Patole· Natural Resources for Human...· 0 citations
Phoneme-based multilingual automatic speech recognition (ASR) can share acoustic evidence across languages more directly than language-specific subword modeling. When tonal and non-tonal languages are jointly trained, however, their supervision granularity does not match: tonal languages annotate tone-marked vowels, wh...
Saierdaer Yusuyin, Nanling Jiang, Hao Huang et al.· 0 citations
Pronunciation assessment requires acoustic evidence that is temporally precise, diagnostically meaningful, and faithful to the learner's actual production. However, existing acoustic models often struggle to provide recognition and segmentation evidence simultaneously. CTC-based phone recognizers can predict phone sequ...
Hao-Peng Geng, Jiun-Ting Li, Daisuke Saito et al.· 0 citations
Introduction Pediatric speech sound disorders (SSDs) affect many young children and are commonly assessed through auditory-perceptual judgments and IPA transcription, which are limited by listener bias, variable interrater reliability, and difficulty attributing deviations at the phoneme level. We present a development...
Chethana Saligram, Vishal Shrivastava, M. Speights· Frontiers in Human Neuroscie...· 0 citations
We study speaker-disjoint accent identification for L2 English, where the goal is to predict a speaker's first-language (L1) background from English pronunciation. Most existing systems classify accents using a single utterance-level representation, but such global representations can obscure pronunciation cues that de...
Yang-Yang Qu, Massimiliano Todisco, Nicholas W. D. Evans· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.