Aug 2026· NED University Journal of Research· 0 citations· 15 references
TL;DR
The Multi-Frame Transformer model can effectively predict intended words from lip movement patterns, offering a foundation for future communication aids for people with aphasia and show that motor signals are more critical for correct prediction.
Abstract
Aphasia affects a person's ability to retrieve words and coordinate the movements of the speech articulators, resulting in impaired speech production and pronunciation. Current word prediction tools do not use lip movement patterns, which limits their usefulness for people with speech disorders. This study aims to predict the word a person intends to say by analyzing both lexical activation and articulatory stability. A Multi-Frame Transformer (MFT) model that processes five consecutive frames of lip movements is developed, along with a proposed Single-Frame Transformer (SFT) and three pre-trained comparative models including BERT, RoBERTa, and a 3D CNN. The proposed MFT model achieved an 84.50% accuracy, outperforming BERT with 63.33%, the 3D CNN with 51.11%, the Single-Frame Transformer with 26.67%, and RoBERTa with11.11%. The proposed model maintained 74% accuracy when word-finding information was missing and 58% accuracy when lip movement information was missing. These findings show that motor signals are more critical for correct prediction. The Multi-Frame Transformer (MFT) model can effectively predict intended words from lip movement patterns, offering a foundation for future communication aids for people with aphasia.
Speech therapists often face difficulties diagnosing impairments due to the lack of efficient tools for transcribing speech into the International Phonetic Alphabet (IPA). This work addresses this challenge with Broca, a Conformer-based deep learning system pretrained on 8 days of adult speech and fine-tuned on a 165-m...
N. Barbaro, Cristina Gena, F. Petriglia et al.· 0 citations
Monitoring cognitive impairment (CI) in motor neuron disease (MND) is essential for timely treatment and care, yet challenging due to co-occurring speech difficulties. The Edinburgh Cognitive and Behavioural ALS Screen (ECAS) provides a robust metric for CI assessment, with the Verbal Fluency Index (VFI) a central elem...
B. Mirheidari, Leslie Ing, Daniel Blackburn et al.· 0 citations
In English pronunciation teaching, prosody has long received less attention than segmental phoneme training, although stress, rhythm, and intonation strongly affect speech intelligibility and naturalness. Existing technological aids mainly focus on phoneme-level correction and still lack fine-grained diagnosis and adap...
LoRA-MoE, a parameter-efficient adaptation method that combines Low-Rank Adaptation with a mixture of experts (MoE) to improve speech recognition for individuals with Parkinson’s disease (PD), demonstrates consistent improvements and stable performance across all severity levels, and its performance is robust to the nu...
Seojin Yoon, Ryul Kim, Sang-Min Lee· IEEE Access· 0 citations
Speech analysis has potential as a non-invasive screening modality for Parkinson’s disease (PD), a neurodegenerative disorder that can affect motor coordination and communication. Conventional speech-based approaches often rely on handcrafted acoustic descriptors or isolated machine-learning models and may not jointly...
V. V, A. Dumka· VFAST Transactions on Softwa...· 0 citations
This work proposes an acoustics-based assessment approach that employs a convolutional neural network to extract rhythm features directly from the speech amplitude envelope, motivated by evidence linking low-frequency modulations to rhythm perception.
João Lima, Lucas H. Ueda, P. Costa· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.