Abstract Artificial intelligence-based approaches to speech analysis have the potential to assist with objective speech error analysis in aphasia but off-the shelf tools often fail to detect speech errors due to prioritizing ‘fluent transcription’. Speech production errors (dysfluencies) are hallmark diagnostic features of the nonfluent and logopenic variants of primary progressive aphasia, yet they can be challenging to detect and characterize even by expert clinicians. This study aimed to evaluate whether the novel automated lightweight Scalable Speech Dysfluency Modeling system, specifically designed to detect dysfluencies, could accurately distinguish primary progressive aphasia variants using voice recordings of individuals reading a brief passage. Participants included a total of 104 individuals, 40 with non-fluent primary progressive aphasia and 40 with logopenic primary progressive aphasia (matched on disease severity), and 24 healthy controls who read aloud the ‘Grandfather Passage’ as part of a widely used motor speech evaluation. We automatically extracted ten speech error (dysfluency) variables, including insertions, replacements, and deletions at both phoneme- and word-levels, and phoneme-level prolongations and repetitions. Group differences were assessed via ANOVAs controlling for age, education, and disease severity (Mini-Mental State Examination and Clinical Dementia Rating sum-of-boxes). To test clinical relevance, we performed correlation analyses with motor speech evaluation ratings provided by experienced speech-language pathologists (i.e. gold standard) within the non-fluent primary progressive aphasia group. Classification performance was assessed by training random forest and XGBoost machine-learning models including 5-fold cross-validation. All individuals read the entire passage in less than five minutes. The lightweight Scalable Speech Dysfluency Modeling system detected 8 of the 10 predefined dysfluency features at sufficient frequency to include them in subsequent analyses. All eight features distinguished primary progressive aphasia from controls (P < 0.006). Individuals with non-fluent primary progressive aphasia made more errors than logopenic primary progressive aphasia on every feature (all P < 0.023). Each feature showed a moderate positive correlation with a global motor speech evaluation apraxia/dysarthria score (r = 0.31–0.56; P < 0.001–0.053). Together, the eight features were able to classify non-fluent versus logopenic at area under the receiver operating characteristic curve = 0.798 (held-out random forest), 0.699 (held-out XGBoost), 0.714 (cross-validated random forest), and 0.704 (cross-validated XGBoost). In sum, automated speech error analysis accurately distinguished non-fluent and logopenic variants using a brief reading task. This quick error-sensitive scalable artificial intelligence system has the potential of providing a practical tool to aid diagnosis in aphasia and motor speech disorders.
Discourse analysis can reliably predict real-world communicative success, yet is rarely implemented clinically, due to challenges in manually generating stimulus-specific main concept inventories (MCIs) and scoring patient narratives. We evaluated whether automatically derived macrolinguistic (main concept, sentiment)...
M. J. Marte, S. Lee, S. Wang et al.· medRxiv· 0 citations
This work analyzed brief story retellings from 86 patients with left-hemisphere stroke and derived discrete linguistic features and embeddings with Large Language Models, providing proof of concept for a fast, largely automated discourse screener of acute LI.
Monitoring cognitive impairment (CI) in motor neuron disease (MND) is essential for timely treatment and care, yet challenging due to co-occurring speech difficulties. The Edinburgh Cognitive and Behavioural ALS Screen (ECAS) provides a robust metric for CI assessment, with the Verbal Fluency Index (VFI) a central elem...
B. Mirheidari, Leslie Ing, Daniel Blackburn et al.· 0 citations
Speech therapists often face difficulties diagnosing impairments due to the lack of efficient tools for transcribing speech into the International Phonetic Alphabet (IPA). This work addresses this challenge with Broca, a Conformer-based deep learning system pretrained on 8 days of adult speech and fine-tuned on a 165-m...
N. Barbaro, Cristina Gena, F. Petriglia et al.· 0 citations
This work systematically evaluates edge-oriented ASR-LLM pipelines for individuals with language impairments using comparison studies and ablation experiments across aphasia, child language impairment, and dementia datasets to identify transcript errors, repetition, noise, and input length as factors affecting system p...
Automatic dysarthric speech detection approaches can support traditional clinical diagnosis, which relies on costly and time-consuming evaluation by a speech and language pathologist. Existing automatic approaches predominantly rely on deep learning (DL). More recently, Large Audio Language Models (LALMs) have emerged...
Mahdi Amiri, H. O. Shahreza, Pascal Frossard et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.