Skip to content
Open access

Shared latent representations of speech production for cross-patient speech decoding

Jul 2026 · Nature Communications · Vol 17 · 0 citations · 54 references
Medicine

TL;DR

By aligning patient-specific neural data to a shared latent space, it is shown that speech BCIs can be trained on data combined across patients, enabling cross-patient speech BCIs and support future speech BCIs that are more accurate and rapidly deployable.

Abstract

Speech brain-computer interfaces (BCIs) can restore communication in individuals with neuromotor disorders who are unable to speak. However, current speech BCIs limit patient usability and successful deployment by requiring large volumes of patient-specific data collected over long periods of time. A promising solution to facilitate usability and accelerate their successful deployment is to combine data from multiple patients. This has proven difficult, however, due to differences in user neuroanatomy, varied placement of electrode arrays, and sparse sampling of targeted anatomy. Here, by aligning patient-specific neural data to a shared latent space, we show that speech BCIs can be trained on data combined across patients. Using canonical correlation analysis and high-density micro-electrocorticography (μECoG), we uncovered shared neural latent dynamics with preserved micro-scale speech information. This approach enabled cross-patient decoding models to achieve improved performance relative to patient-specific models facilitated by the high resolution and broad coverage of μECoG. Our findings support future speech BCIs that are more accurate and rapidly deployable, ultimately improving the quality of life for people with impaired communication from neuromotor disorders. Current speech brain-computer interfaces (BCIs) rely on patient-specific decoding approaches. Here, the authors show that patient-specific data can be aligned to a shared space that preserves speech information, enabling cross-patient speech BCIs.

Read PDF

Similar papers

Open access Jul 2026

A generalizable speech neuroprosthesis

A transformer-based decoder model trained jointly across six intracortical speech BCI participants reveals how to pool intracortical data across people to yield more accurate, generalizable, and rapidly-deployable decoding models.

Zachery M. Fogg, N. Card, M. Wairagkar et al. · 0 citations
Preprint Aug 2026

Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings

Decoding perceived speech from non-invasive brain recordings has garnered significant attention in recent years due to its wide range of potential applications. However, existing methods face considerable challenges in cross-subject decoding, primarily due to limited generalizability and the absence of explicit mechanisms for extracting subject-consistent information. These limitations result in high training costs and suboptimal decoding performance. To address these challenges, we propose an innovative Cross-Subject Perceived Speech Decoding (CPSD) framework, which comprises two training stages: source model pre-training and personal specialization. In the source model pre-training stage, contrastive learning is employed to capture shared representations across multiple source subjects. Subsequently, personal specialization initializes the model for the target subject by extracting consistent components from the source model and fine-tuning it using target subject data. Additionally, we introduce the Positional Encoding-based Spatial Attention (PESA) module, which remaps MEG/EEG data into a standardized reference space, thereby enhancing cross-subject consistency and facilitating model training. We evaluate the proposed CPSD framework on three perceived speech neural datasets encompassing different modalities and languages. The results demonstrate that our framework outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top-10 accuracy on the Armeni 2022, PKUEEG 2025, and Broderick 2018 datasets, respectively. Further analyses confirm the effectiveness, efficiency, and robustness of the proposed approach.

Aoke Zhang, Bo Wang, Xihong Wu et al. · 0 citations
Preprint Aug 2026

Motor, Cognitive, or Corpus? What Survives Cross-Lingual Transfer in Speech-Based Parkinsons Disease Detection

Self-supervised learning (SSL) speech representations achieve strong performance for Parkinson's disease (PD) detection within individual corpora. However, it remains unclear whether these models capture disease-related characteristics or exploit dataset-specific confounds, particularly since most SSL backbones are pretrained exclusively on healthy speech. To investigate this question, we perform a layer-wise analysis of nine SSL speech backbones using a low-capacity logistic regression probe across three languages. We structure the evaluation as multiple scenarios that progressively introduce distribution shifts in participant identity, recording conditions, language, and pathology. Our results reveal two key findings. First, layer selection is highly corpus-dependent: the optimal representation layer is determined primarily by the source dataset rather than by the SSL architecture itself. Second, the transferred discriminative signal lacks pathological specificity: classifiers trained to detect PD assign similarly high probabilities to both PD and dementia speech in the target corpus. These results highlight critical limitations that must be addressed before speech-based pathology recognition models can be reliably deployed in clinical settings.

S. Kopar, Sam Gijsen, Abner Hernandez et al. · 0 citations
Open access Jul 2026

Decoding Chinese speech across multiple neural conditions via EEG: dataset construction and interpretability driven spatial optimization

The integration of artificial intelligence (AI) and brain-computer interfaces (BCIs) technologies shows great potential in assisting patients with speech impairments and improving cognitive-linguistic decline. Electroencephalogram (EEG) based BCIs, characterized by non-invasiveness, low cost, and high temporal resolution, hold significant application value in speech decoding and cognitive rehabilitation. Currently, most mainstream public EEG datasets rely on Western languages. As a tonal language, Chinese Mandarin differs significantly from Western languages in speech production mechanisms, making existing data insufficient to support future BCI research for Mandarin-speaking patients. To address this gap, we establish a systematic Mandarin EEG dataset and conduct effective speech decoding and related analyses. We design four distinct experimental conditions, namely overt, overt-noisy, intend, and imagine, to simulate different types of speech disorders in clinical scenarios. Using typical Mandarin tonal-vowels and common vocabularies as stimuli, we construct an EEG dataset collected from a healthy adult. We evaluate the speech decoding performance using short-time Fourier transform combined with support vector machine (STFT-SVM) and EEG-Conformer models. Furthermore, we design a multi-task architecture based on the EEG-Conformer to perform a unified decoding task for the two stimulus types and a classification task across the four dataset conditions. To interpret the model, we combine Shapley value computation and decision trees to calculate the importance of different electrodes during classification. Experimental results show that the models achieve effective decoding on our dataset. The EEG-Conformer model performs significantly above chance level across all data, reaching an accuracy of 69.83% in normal speaking conditions and up to 61.46% in conditions simulating speech disorders. In the multi-task setting, the classification accuracy across different conditions exceeds 97%. By utilizing the important electrodes identified through interpretability methods as new feature inputs, the classification performance further improves even with a reduction of over 50% in the channels. These results demonstrate the potential of neural signal decoding technologies in communication assistance, reveal the decodability of Chinese Mandarin EEG datasets, and provide feasible recommendations for the future design of Chinese BCI applications.

Haoming Wang, Gaoyuan Zhang, Xurong Xie et al. · 0 citations
Preprint Aug 2026

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

Paired MEG occlusion shows that 15 of 19 stimulus features contribute, with the largest effects for silence, sound intensity, vowels, and acoustic onsets, indicating that activity without narrative structure carries less recoverable information than activity during coherent speech.

Ilia Semenkov, Daria Kleeva, I. Dakhtin et al. · 0 citations
Preprint Aug 2026

Decoding silent reading from non-invasive EEG

Non-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person's spontaneous inner monologue cannot be collected, and the available proxy paradigms (cued repetitive and retrospectively reported generative inner speech) are slow to acquire, poorly time-locked, and subject compliance is unverifiable. We therefore treat silent reading as a scalable proxy task and ask how much lexical and semantic information a contrastive decoder can extract from it. We report an open-vocabulary analysis of approximately 240,000 word presentations recorded from a single densely-sampled participant across 393 runs (ca. 49 h) of 19-channel dry-electrode EEG. Words from continuous narrative text were presented in rapid serial visual presentation, with typography randomised on every trial to partially decorrelate word identity from low-level visual form. A convolutional EEG encoder, optionally followed by a causal transformer, was trained with a CLIP-style contrastive objective to align short EEG windows with hidden-state embeddings of the presented word taken from a large language model. Decoding, evaluated as word-grouped top-10 retrieval against permutation baselines, was reliably above chance, extended to mid-frequency and rare words, and scaled log-linearly with training-data volume with no sign of saturation. Removing occipital and posterior-temporal electrodes reduced the word-level gain by roughly one third but left context tracking unchanged. Control analyses separate word-level decoding from narrative context tracking and from a non-neural positional prior introduced by the transformer's positional embedding. These results establish that open-vocabulary word-level information is recoverable from EEG during silent reading, and that decoding is data-limited rather than saturated.

I. Marquardt, A. Alchanat, Priyanka Jain · 0 citations