A transformer-based decoder model trained jointly across six intracortical speech BCI participants reveals how to pool intracortical data across people to yield more accurate, generalizable, and rapidly-deployable decoding models.
By aligning patient-specific neural data to a shared latent space, it is shown that speech BCIs can be trained on data combined across patients, enabling cross-patient speech BCIs and support future speech BCIs that are more accurate and rapidly deployable.
Z. Spalding, S. Duraivel, S. Rahimpour et al.· Nature Communications· 0 citations
Brain-computer interfaces (BCIs) offer a promising solution to speech loss due to neurological injury by decoding intended speech directly from brain activity. While recent BCIs have restored high-accuracy text-based communication, they fail to provide instantaneous voice output essential for the natural flow of conversation. Brain-to-voice BCIs address this gap by decoding voice directly from neural signals. However, even the state-of-the-art (SOTA) BCI-synthesized voice is not yet intelligible enough for real-world adoption. We introduce brain2voice 2.0, a new multimodal Transformer-based BCI decoder architecture capable of synthesizing highly intelligible voice from intracortical neural signals in real-time. Brain2voice 2.0 is trained on continuous and custom-tokenized acoustic targets and phoneme targets, leveraging their complementary speech information. We use self-supervised and adversarial training objectives that enhance acoustic feature quality and improve synthesis intelligibility. At each 10 ms timestep, the model causally outputs continuous and tokenized acoustic features for real-time voice synthesis as well as time-aligned phoneme predictions (raw phoneme error rate: 7%, comparable to the latest brain-to-text models). We evaluated this new approach on our prior intracortical brain-to-voice benchmark dataset (Wairagkar et al. 2025). Naive human listeners transcribed brain2voice 2.0 synthesized voice with a word error rate of 5.24%—an 8× improvement in intelligibility over previous SOTA results (43.75%). Brain2voice 2.0 demonstrates that highly intelligible real-time voice synthesis from neural signals is achievable, for the first time crossing the intelligibility threshold necessary for clinically viable brain-to-voice BCIs for people with paralysis.
M. Wairagkar, Aparna Srinivasan, N. Card et al.· bioRxiv· 0 citations
State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent decoders, and how the output target (phonetic vs.\ character) interacts with that choice. On the public Brain-to-Text'25 benchmark, we study a controlled 2x2 grid (GRU vs.\ hybrid Mamba decoder; phonetic vs.\ character targets) trained with a CTC objective under one reproducible protocol. The recurrent baseline remains strongest: the best phonetic GRU reaches 12.62\% PER and 21.19\% WER, while the best textual GRU after LM rescoring reaches 13.39\% CER and 26.28\% WER. The Mamba hybrid is competitive but does not surpass it. Ablations isolate architectural contributions, and error analysis shows representation-dependent failures: articulatory-like phoneme confusions vs.\ lexical and word-boundary errors.
Lucas Zamora Vera, Jose A. Gonzalez-Lopez· 0 citations
Non-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person's spontaneous inner monologue cannot be collected, and the available proxy paradigms (cued repetitive and retrospectively reported generative inner speech) are slow to acquire, poorly time-locked, and subject compliance is unverifiable. We therefore treat silent reading as a scalable proxy task and ask how much lexical and semantic information a contrastive decoder can extract from it. We report an open-vocabulary analysis of approximately 240,000 word presentations recorded from a single densely-sampled participant across 393 runs (ca. 49 h) of 19-channel dry-electrode EEG. Words from continuous narrative text were presented in rapid serial visual presentation, with typography randomised on every trial to partially decorrelate word identity from low-level visual form. A convolutional EEG encoder, optionally followed by a causal transformer, was trained with a CLIP-style contrastive objective to align short EEG windows with hidden-state embeddings of the presented word taken from a large language model. Decoding, evaluated as word-grouped top-10 retrieval against permutation baselines, was reliably above chance, extended to mid-frequency and rare words, and scaled log-linearly with training-data volume with no sign of saturation. Removing occipital and posterior-temporal electrodes reduced the word-level gain by roughly one third but left context tracking unchanged. Control analyses separate word-level decoding from narrative context tracking and from a non-neural positional prior introduced by the transformer's positional embedding. These results establish that open-vocabulary word-level information is recoverable from EEG during silent reading, and that decoding is data-limited rather than saturated.
I. Marquardt, A. Alchanat, Priyanka Jain· 0 citations
Imagined speech refers to the internal rehearsal of speech without articulation. Decoding and classifying imagined speech assists motor nerve disabled patients to communicate their needs to their caretakers. ElectroEnchephoGraphy (EEG) signals of imagined speech can be decoded into target commands. Effective decoding requires localization of electrodes to isolate neural signals pertinent to imagined speech. In this research, imagined speech signals of 10 healthy individuals for eight utilitarian words were extracted using 21-channel EEG acquisition device. Time domain features were extracted and analyzed using Extra tree classifier (ETC) in subject-specific manner. Using the Gini index, the eight most important spatial features for imagined speech were isolated. Frequency domain features across five bands of brain waves were analyzed from these isolated spatial positions using Fast Fourier transform (FFT). Principal Component Analysis (PCA) was employed for dimensionality reduction and classification was done using ETC, Decision Tree (DT) and KNN. ETC performed well, with a mean accuracy of 88.77%. To improve classification performance, Long Short Term Memory (LSTM) model with a Sliding Window and Attention layer (LSTM-SWA) was implemented. LSTM-SWA achieved a mean accuracy of 92.02%, as it offers the advantage of interpretability by identifying relevant temporal segments of neural activity associated with imagined speech. These findings demonstrate the need to identify effective electrode positions and frequency domain features to design an Alternative and Augmentative Communication (AAC) device using imagined speech signals.
K. Vaishnavi, G. S. Sadasivam· Scientific Reports· 0 citations
Abstract Comprehending connected speech is critical for human interaction and is vulnerable in post-stroke aphasia. Understanding the neural mechanisms underlying impaired speech listening is necessary for accurate and effective assessment and treatment. Neural speech tracking methods offer a window into naturalistic speech processing and may reveal causal contributions to comprehension. EEG was recorded during story listening in 15 people with aphasia with left temporal lesions (temporal group), 14 people with aphasia with left frontal lesions (frontal group) and 15 age- and hearing-matched controls (control group). All participants listened to clear and unintelligible stories (∼12 min per condition). Controls additionally listened to low-intelligibility stories that equated comprehension success to the temporal group. Envelope tracking was measured at syllable- (theta) and multi-syllable- (delta) rates. Neural decoding and encoding analyses measured global (whole-brain) and local (sensor-wise) speech tracking, respectively. Group comparisons used linear mixed-effects regression. Linear and quadratic relationships assessed associations between neural tracking and behavioural measures of comprehension while accounting for covariates. Behavioural measures of comprehension showed the temporal group were significantly impaired compared to both frontal and control groups. Comparison of intelligible and unintelligible speech found that intelligibility affected delta but not theta tracking—suggesting that delta tracking is linked to higher-order linguistic processing and theta tracking reflects sensory responses. Theta and delta envelope decoding reflected group/comprehension status: tracking was reduced in the temporal group compared to the control group, but tracking was not significantly different when control comprehension was behaviour-matched via speech degradation. Delta encoding produced a similarly behaviour-linked pattern, but theta encoding was lesion—rather than behaviour-sensitive, in that reduced tracking was observed in temporo-parietal sensors in both aphasia groups. Theta tracking correlated with comprehension in a lesion-specific manner. Positive correlations between theta tracking and comprehension were found in the temporal group and negative correlations in the frontal and control groups. This produced a quadratic (inverted-U) relationship between theta tracking and comprehension success across all participants. The control-like, negative theta tracking–comprehension correlations in the frontal group are consistent with listening effort effects in which greater task difficulty result in greater tracking. In the temporal group, where comprehension was impaired, better envelope tracking at syllable rates may help build a stable representation of the speech stream supporting phonological and lexical analysis and comprehension. These results indicate that envelope tracking reflects a combination of lesion and symptom profiles and is a viable method for investigating the mechanisms of speech comprehension in aphasia.
Guangting Mai, Emily Upton, Timothy D. Griffiths et al.· Brain Communications· 0 citations