Skip to content

Author

Xurong Xie

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception

Decoding speech information directly from scalp electroencephalography (EEG) into text provides a potential non-invasive neural communication pathway for individuals with severe speech and motor impairments. Compared with invasive approaches such as electrocorticography, EEG is safer and more widely deployable, yet substantially more challenging to decode.This challenge is exacerbated for Chinese sentence decoding, which must handle a high-dimensional output space with thousands of characters, severe inter-subject variability, and low signal-to-noise ratios for text alignment.Existing methods commit to a single supervisory axis---either text semantics or audio acoustic features---yet neither can simultaneously satisfy the demands of sentence-level discriminability and fine-grained temporal resolution required for large-vocabulary Chinese decoding. We introduce EEGAlign, a novel parameter-efficient framework that jointly aligns EEG with two axes---text alignment with BGE-M3 text embeddings and audio alignment with wav2vec~2.0 speech features via contrastive learning followed by CTC character-sequence decoding. On ChineseEEG-2 data, EEGAlign yields state-of-the-art closed-set sentence classification performance, reaching up to 82.37% Top-1 accuracy on Reading Aloud EEG and 41.43% on Passive Listening EEG out of 101 candidates. Ablation studies show that the two alignment axes are highly complementary: combining them yields consistently better performance than either alone. To the best of our knowledge, this is the first study on decoding large-vocabulary Chinese sentences from non-invasive EEG during overt speech production, and achieving strong classification performance with relatively large closed-set candidate-sentence setting.

Tian Zheng, Xurong Xie, Xinxin Zhu et al. · 0 citations
Open access Jul 2026

Decoding Chinese speech across multiple neural conditions via EEG: dataset construction and interpretability driven spatial optimization

The integration of artificial intelligence (AI) and brain-computer interfaces (BCIs) technologies shows great potential in assisting patients with speech impairments and improving cognitive-linguistic decline. Electroencephalogram (EEG) based BCIs, characterized by non-invasiveness, low cost, and high temporal resolution, hold significant application value in speech decoding and cognitive rehabilitation. Currently, most mainstream public EEG datasets rely on Western languages. As a tonal language, Chinese Mandarin differs significantly from Western languages in speech production mechanisms, making existing data insufficient to support future BCI research for Mandarin-speaking patients. To address this gap, we establish a systematic Mandarin EEG dataset and conduct effective speech decoding and related analyses. We design four distinct experimental conditions, namely overt, overt-noisy, intend, and imagine, to simulate different types of speech disorders in clinical scenarios. Using typical Mandarin tonal-vowels and common vocabularies as stimuli, we construct an EEG dataset collected from a healthy adult. We evaluate the speech decoding performance using short-time Fourier transform combined with support vector machine (STFT-SVM) and EEG-Conformer models. Furthermore, we design a multi-task architecture based on the EEG-Conformer to perform a unified decoding task for the two stimulus types and a classification task across the four dataset conditions. To interpret the model, we combine Shapley value computation and decision trees to calculate the importance of different electrodes during classification. Experimental results show that the models achieve effective decoding on our dataset. The EEG-Conformer model performs significantly above chance level across all data, reaching an accuracy of 69.83% in normal speaking conditions and up to 61.46% in conditions simulating speech disorders. In the multi-task setting, the classification accuracy across different conditions exceeds 97%. By utilizing the important electrodes identified through interpretability methods as new feature inputs, the classification performance further improves even with a reduction of over 50% in the channels. These results demonstrate the potential of neural signal decoding technologies in communication assistance, reveal the decodability of Chinese Mandarin EEG datasets, and provide feasible recommendations for the future design of Chinese BCI applications.

Haoming Wang, Gaoyuan Zhang, Xurong Xie et al. · 0 citations