Skip to content

MULTI-FRAME TRANSFORMER-BASED COMMUNICATION SIGNAL MODELING FOR WORD PREDICTION IN APHASIA

Aug 2026 · NED University Journal of Research · 0 citations · 15 references

TL;DR

The Multi-Frame Transformer model can effectively predict intended words from lip movement patterns, offering a foundation for future communication aids for people with aphasia and show that motor signals are more critical for correct prediction.

Abstract

Aphasia affects a person's ability to retrieve words and coordinate the movements of the speech articulators, resulting in impaired speech production and pronunciation. Current word prediction tools do not use lip movement patterns, which limits their usefulness for people with speech disorders. This study aims to predict the word a person intends to say by analyzing both lexical activation and articulatory stability. A Multi-Frame Transformer (MFT) model that processes five consecutive frames of lip movements is developed, along with a proposed Single-Frame Transformer (SFT) and three pre-trained comparative models including BERT, RoBERTa, and a 3D CNN. The proposed MFT model achieved an 84.50% accuracy, outperforming BERT with 63.33%, the 3D CNN with 51.11%, the Single-Frame Transformer with 26.67%, and RoBERTa with11.11%. The proposed model maintained 74% accuracy when word-finding information was missing and 58% accuracy when lip movement information was missing. These findings show that motor signals are more critical for correct prediction. The Multi-Frame Transformer (MFT) model can effectively predict intended words from lip movement patterns, offering a foundation for future communication aids for people with aphasia.

View source

Similar papers

#machine learning Preprint Sep 2026

Deep Learning Techniques for Phoneme Recognition in Italian Children's Speech

Speech therapists often face difficulties diagnosing impairments due to the lack of efficient tools for transcribing speech into the International Phonetic Alphabet (IPA). This work addresses this challenge with Broca, a Conformer-based deep learning system pretrained on 8 days of adult speech and fine-tuned on a 165-m...

N. Barbaro, Cristina Gena, F. Petriglia et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Automatic estimation of verbal fluency index in people with Motor Neuron Disease using ASR alignment and pause modelling

Monitoring cognitive impairment (CI) in motor neuron disease (MND) is essential for timely treatment and care, yet challenging due to co-occurring speech difficulties. The Edinburgh Cognitive and Behavioural ALS Screen (ECAS) provides a robust metric for CI assessment, with the Verbal Fluency Index (VFI) a central elem...

B. Mirheidari, Leslie Ing, Daniel Blackburn et al. · 0 citations
Open access Aug 2026

Precise Guidance for English Prosody Perception Pronunciation Based on Deep Learning

In English pronunciation teaching, prosody has long received less attention than segmental phoneme training, although stress, rhythm, and intonation strongly affect speech intelligibility and naturalness. Existing technological aids mainly focus on phoneme-level correction and still lack fine-grained diagnosis and adap...

Tingting Liu, Qing Li · 0 citations
Open access 2026

LoRA-MoE Fine-Tuning for Improved Speech Recognition in People With Parkinson’s Disease

LoRA-MoE, a parameter-efficient adaptation method that combines Low-Rank Adaptation with a mixture of experts (MoE) to improve speech recognition for individuals with Parkinson’s disease (PD), demonstrates consistent improvements and stable performance across all severity levels, and its performance is robust to the nu...

Seojin Yoon, Ryul Kim, Sang-Min Lee · 0 citations
Review Open access Sep 2026

VoiceFusionNet: A Hybrid CNN--Transformer Framework for Speech-Based Parkinson’s Disease Screening Using Speech Signal Analysis

Speech analysis has potential as a non-invasive screening modality for Parkinson’s disease (PD), a neurodegenerative disorder that can affect motor coordination and communication. Conventional speech-based approaches often rely on handcrafted acoustic descriptors or isolated machine-learning models and may not jointly...

V. V, A. Dumka · 0 citations
Preprint Sep 2026

Automated Assessment of L2 Speech Rhythm Using Low-Frequency Amplitude Modulations

This work proposes an acoustics-based assessment approach that employs a convolutional neural network to extract rhythm features directly from the speech amplitude envelope, motivated by evidence linking low-frequency modulations to rhythm perception.

João Lima, Lucas H. Ueda, P. Costa · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.