Leveraging automatic forced alignment of reading passages in dysarthria
Abstract
Manual acoustic-phonetic segmentation of dysarthric speech is challenging due to the degraded nature of the acoustic signal. The Montreal Forced Aligner (MFA) is a gold-standard approach for automatic segmentation in non-disordered speech, but its efficacy for speakers with dysarthria has not been evaluated. We evaluate the MFA’s performance on vowel segmentation in a standardized reading passage produced by five talkers with dysarthria and five healthy controls. MFA input includes speech audio and an orthographic transcript and uses pre-trained acoustic models and pronunciation dictionaries to time-align word and phone boundaries. We manipulate three alignment conditions that differ in the amount of pre-processing a researcher may choose to perform on their input transcripts, namely (1) the full text of the reading passage, (2) the passage segmented into utterances, and (3) the passage segmented into utterances, with speech errors corrected in the text. We compare the force aligned output with manually segmented vowels as a function of condition, speaker group, and vowel. Alignment across all conditions was high ( > 90% accuracy) for controls, but varied widely for dysarthric speech. Alignment accuracy for dysarthric speakers ranged from 20% (full-text) to 76.4% (segmented-corrected-text). Findings will guide best-practices for leveraging automatic acoustic tools in disordered speech research.