Skip to content
Preprint

EEG2MOTION: Towards Open-Vocabulary Human Motion Synthesis from Non-invasive Brain Signals

Aug 2026 · 0 citations · 33 references
Engineering

TL;DR

This work proposes EEG-conditioned Masked Motion Model (EMMM), a generative framework that unites an EEG encoder with a motion decoder to synthesize continuous, full-body human motions directly from brain activity, opening a new direction toward generative and open-vocabulary motor BCIs.

Abstract

Human motion is governed by a hierarchical motor system where the brain provides high-level intentions and lower-level structures coordinate detailed dynamics. Existing brain-computer interfaces (BCIs) typically oversimplify this into constrained classification or low-dimensional control, failing to capture the richness of natural movement. Bridging this gap to achieve open-vocabulary, full-body motion synthesis remains challenging due to the substantial cross-modal divergence between sparse neural signals and high-dimensional kinematics, as well as the lack of large-scale paired EEG-motion datasets. To address this, we introduce EEG2MOTION, the first EEG-motion-text dataset for human motion synthesis, comprising nearly 20,000 paired samples across thousands of motions. Using this dataset, we first demonstrate via multimodal contrastive learning that non-invasive EEG embeddings can be effectively aligned with text, video, and motion representations to decode high-level semantics. We then propose EEG-conditioned Masked Motion Model (EMMM), a generative framework that unites an EEG encoder with a motion decoder to synthesize continuous, full-body human motions directly from brain activity. Experimental results show that EMMM generates coherent and realistic motion sequences from non-invasive brain signals. To the best of our knowledge, this is the first work to generate diverse full-body human motions from non-invasive brain signals, opening a new direction toward generative and open-vocabulary motor BCIs. See our project page: https://yulom.github.io/EEG2MOTIONdemopage/.

View source

Similar papers

Open access Oct 2025

Cortical-SSM: a deep state space model for motor imagery decoding from EEG signals

Objective. Classification of electroencephalogram (EEG) signals obtained during motor imagery (MI) has substantial application potential, including for communication assistance and rehabilitation support for patients with motor impairments. These signals remain inherently susceptible to physiological artifacts (e.g. eye blinking, swallowing), which pose persistent challenges. Although Transformer-based approaches for classifying EEG signals have been widely adopted, they often struggle to capture fine-grained dependencies within them. Approach. To overcome these limitations, we propose Cortical-SSM, a novel architecture that extends deep state space models to capture integrated dependencies of EEG signals across temporal, spatial, and frequency domains. We validated our method across two large-scale public MI EEG datasets containing more than 50 subjects. Main results. Our method outperformed baseline methods on the two benchmarks. Furthermore, visual explanations derived from our model indicate that it effectively captures neurophysiologically relevant regions of EEG signals. Significance. These results indicate that Cortical-SSM provides a robust and interpretable alternative to attention-based architectures for MI EEG decoding. By enabling physiologically grounded feature learning, our method advances the reliability of subject-independent EEG classification and supports the development of practical and clinically deployable brain–computer interface systems.

Shuntaro Suzuki, Shunya Nagashima, Komei Sugiura · 0 citations
Preprint Jul 2026

EEGForceFusion: Joint Tokenised-Continuous Representation Learning for Subject-Independent Grasp Force Decoding

Brain-machine interfaces provide a link between neural activity and external devices, enabling restoration of motor function and advancing human-machine interaction using non-invasive electroencephalography (EEG). However, continuous grasp force decoding remains challenging due to complex temporal dynamics, high inter-subject variability, and limited generalisation of existing approaches. To address this, we propose a hybrid EEG decoding framework that jointly models continuous and tokenised representations, enabling capture of both fine-grained neural structure and long-range temporal dependencies. The proposed approach integrates convolutional-recurrent representation learning, quantisation-based tokenisation, and transformer-based temporal modelling within a unified fusion-based regression architecture. Experimental evaluation on the WAY-EEG-GAL dataset under strict leave-one-subject-out conditions achieves $R^2$ = 0.817 in offline settings and $R^2$ = 0.793 in simulated real-time evaluation, with latency suitable for real-time deployment. These results demonstrate strong cross-subject generalisation and highlight the practicality of hybrid continuous-tokenised representations for real-time EEG-based force decoding in assistive robotics, neuro-rehabilitation, and human-machine interaction.

Sankalp Sunil Turankar, Y. Meena · 0 citations
Book Open access Aug 2026

Neuro-SPO: Physically-Grounded and Affect-Aligned EEG-to-Keyword Decoding

Decoding semantic information from non-invasive EEG remains a formidable challenge in Brain-Computer Interfaces due to the low signal-to-noise ratio and the complex biophysics of neural dynamics. While large-scale pre-trained models like Whisper provide powerful generic temporal representations, they lack the domain-specific biophysical constraints essential for EEG. Conversely, pure physical operators capture topology but miss the rich semantic priors. To bridge this gap, we propose the Neural Spectral-Physical Operator (Neuro-SPO), a robust framework that reframes EEG-to-Keyword decoding as a continuous representation learning problem. Neuro-SPO synergizes two complementary paradigms: a Pre-trained Temporal Operator derived from the Whisper encoder to extract high-level temporal semantics, and a novel Neural Holistic Physics Mixer (Neuro-HPM) that injects domain-specific physical constraints via dynamic spectral graph interactions. By fusing these streams, Neuro-SPO ensures that the learned representations are both semantically rich and biophysically plausible. Furthermore, to address the ambiguity of neural signals caused by affective variance, we introduce a Structure-Regularized Optimization strategy, employing the Affective Modulation Alignment and Topology-Preserving Ranking Objective to rectify the latent decision space. Extensive experiments on the ZuCo and ChineseEEG benchmarks demonstrate that Neuro-SPO significantly outperforms state-of-the-art methods in ranking metrics and retrieval accuracy. The code is available at https://github.com/Dray-Xu/Neuro-SPO.

Zihua Xu, C. Chen, Tong Zhang · 0 citations
Review Open access Aug 2026

Dynamic Cognitive Prior Generation for Robust EEG-to-Image Decoding

One of the key challenges in brain computer interfaces (BCIs) is to understand the visual perceptual content (VPC) of non-invasive electroencephalography (EEG) signals with high accuracy, without the aid of brain mapping techniques, which is hindered by high inter-subject variations and the non-stationary nature of neural responses. Current methods, including NeuroBridge, use handcrafted, fixed, subject-agnostic transformations of perceptual variance called Cognitive Prior Augmentation (CPA). These static priors, however, have little capability for modelling the dynamic changes of cognition states and individual brain properties, which severely constrain across-subjects generalization. We introduced Dynamic Cognitive Prior Generation (DCPG) a new framework, which can be learned and is prototype based to adaptively generate priors instead of heuristic augmentations. Our approach distills the subject-specific cognitive priors by modelling the attention to a common bank of prototype representations, based on a given EEG trial and based on a learnable subject embedding. Using feature modulation, DCPG can adaptively calibrate the representation of EEG before semantic projection, thus reducing the domain shift between different subjects effectively. It was shown that DCPG can be used to significantly increase the accuracy of inter-subject retrieval on the THINGS-EEG dataset with an accuracy improvement of +3.0% while adding 4.9% more parameters compared to NeuroBridge. This framework is the new state-of-the-art in the field of robust EEG-image decoding, with results in a variety of populations. REFERENCES [1] W. Zhang, S. Wang, Y. Su, X. Li, C. Zhang, and S. Zhong, "Neurobridge: Bio-inspired self-supervised eeg-to-image decoding via cognitive priors and bidirectional semantic alignment," in Proc. AAAI Conf. Artif. Intell., vol. 40, no. 21, pp. 18028-18036, 2026. [2] K. Ahmed, H. A. Mahmoud, H. M. El Hadad, and Y. Afifi, "Generative AI for the reconstruction of visual stimuli from functional magnetic resonance imaging (fMRI) signals," J. Big Data, 2026. [3] H. Yu, Q. Mu, C. Liu, S. Wang, and J. Sun, "Technical system of electroencephalography-based brain–computer interface: Advances, applications, and challenges," Neural Regener. Res., vol. 21, no. 9, pp. 3885-3907, 2026. [4] X. T. Tran et al., "Inter-and Intra-Subject Variability in EEG: A Systematic Survey," arXiv:2602.01019, 2026. [5] S. Shukla et al., "A survey on bridging EEG signals and generative AI: from image and text to beyond," arXiv:2502.12048, 2025. [6] B. Abibullaev, A. Keutayeva, and A. Zollanvari, "Deep learning in EEG-based BCIs: A comprehensive review of transformer models, advantages, challenges, and applications," IEEE Access, vol. 11, pp. 127271-127301, 2023. [7] M. Hemati, C. Rowley, E. Deem, and L. Cattafesta, "De-biasing the dynamic mode decomposition for applied Koopman spectral analysis of noisy datasets," Theor. Comput. Fluid Dyn., vol. 31, no. 4, 2017. [8] Y. Prabhu, "Unveiling Bias in Multimodal Models," Ph.D. dissertation, 2025. [9] L. Guo and W. Zhang, "Geospatial sentiment analytics for hospitality management: predicting investment returns on service attributes across urban micro-zones," J. Quality Assurance Hospitality Tourism, pp. 1-24, 2026. [10] Y. Ke, D. Liang, and K. Shang, "NeuroAlign: Dynamic Dual-Stream Alignment of Perception and Cognition for Zero-Shot Brain-Image Retrieval," in Proc. 2026 Int. Conf. Multimedia Retrieval, pp. 108-117, 2026. [11] A. F. Nia, "Using Advanced Machine Learning Techniques to Recognize Emotional States from a Hybrid fNIRS-EEG System," Ph.D. dissertation, Univ. Auckland, 2025. [12] U. Iqbal, "AI-driven predictive maintenance for US smart manufacturing: Deep learning models for equipment failure prediction and operational resilience," J. Eng. Comput. Intell. Rev., vol. 3, no. 1, pp. 114-138, 2025. [13] S. Shumba, "Multi-Modal Automated Major Depressive Disorder Detection," Ph.D. dissertation, Stellenbosch Univ., 2025. [14] M. Parvan, "ECG-biometrics-bench: A Unified Framework for Reproducible Benchmarking of ECG Biometrics," arXiv:2605.01548, 2026. [15] E. G. Ahsaei, "Constructing and Analyzing Neural Network Dynamics for Information Objectives and Working Memory," Ph.D. dissertation, Washington Univ. St. Louis, 2021. [16] U. Iqbal, "AI-powered supplier risk intelligence: Predicting financial and geopolitical supply chain disruptions in US critical industries," J. Eng. Comput. Intell. Rev., vol. 3, no. 2, pp. 173-193, 2025. [17] P. M. Alamdari, "Multi-Persona Adaptive Recommendation via Cognitive Persona Transitions and Behavioral State Flow Modeling," 2025. [18] R. G. Praveen, P. Cardinal, and E. Granger, "Weakly supervised learning for facial affective behavior analysis: A review," IEEE Trans. Affect. Comput., 2026. [19] U. Imtiaz, F. Amin, and A. Khan, "Crash2Compromise: CrashGuard-FS for Security-Aware Linux File Updates," Multidiscip. Res. Comput. Inf. Syst., vol. 6, no. 2, pp. 294-310, 2026. [20] A. Khan, U. Imtiaz, and F. Amin, "Big Data Analytics in Advanced Cybersecurity: A US Study of Proactive Strategies and Innovative Solutions," Spanish J. Innov. Integrity, vol. 54, pp. 163-180, 2026. [21] S. Hussain, P. K. Jamwal, and P. Van Vliet, "Design synthesis and optimization of a 4-SPS intrinsically compliant parallel wrist rehabilitation robotic orthosis," J. Comput. Des. Eng., vol. 8, no. 6, pp. 1562-1575, 2021. [22] F. Jiang, N. Du, M. Yu, and Q. He, "EmoDNCL+: Dual-stream negative-sample-free contrastive learning with neurophysiological augmentation for EEG emotion recognition," J. King Saud Univ. Comput. Inf. Sci., vol. 38, no. 2, p. 30, 2026. [23] T. Li, Y. Yan, F. Dou, W. Song, and X. Zhang, "Cross-subject generalization for EEG decoding: a survey of deep learning methods," Prog. Biomed. Eng., vol. 8, no. 2, p. 022013, 2026. [24] D. Bailey, "Why Adaptive AI Has No Inner Life: A Structural Account of Function Without Phenomenology," 2026. [25] J. Duan et al., "Critical changes in whole-brain gene networks in response to small-cell lung cancer as revealed by single-nucleus RNA sequencing," Frontiers Immunol., vol. 17, p. 1860628, 2026. [26] U. Iqbal and Y. Bhutto, "Digital transformation through artificial intelligence and advance business analytic in American operational management," J. Theor. Appl. Econometrics, vol. 3, no. 1, pp. 37-50, 2026. [27] U. Iqbal, "AI-enhanced network optimization for electric vehicle charging infrastructure expansion in the United States using graph theory and demand analytics," J. Eng. Comput. Intell. Rev., vol. 2, no. 2, pp. 112-129, 2024. [28] U. Iqbal, S. Bekmez, and F. A. Qurashi, "Operational Risk Management Through Machine Learning and Business Intelligence in U.S. Businesses," Spanish J. Innov. Integrity, vol. 54, pp. 239-253, 2026.

Mehran Ali · 0 citations
Preprint Aug 2026

Decoding silent reading from non-invasive EEG

Non-invasive decoding of inner speech faces a fundamental data problem: a corpus pairing brain activity with a person's spontaneous inner monologue cannot be collected, and the available proxy paradigms (cued repetitive and retrospectively reported generative inner speech) are slow to acquire, poorly time-locked, and subject compliance is unverifiable. We therefore treat silent reading as a scalable proxy task and ask how much lexical and semantic information a contrastive decoder can extract from it. We report an open-vocabulary analysis of approximately 240,000 word presentations recorded from a single densely-sampled participant across 393 runs (ca. 49 h) of 19-channel dry-electrode EEG. Words from continuous narrative text were presented in rapid serial visual presentation, with typography randomised on every trial to partially decorrelate word identity from low-level visual form. A convolutional EEG encoder, optionally followed by a causal transformer, was trained with a CLIP-style contrastive objective to align short EEG windows with hidden-state embeddings of the presented word taken from a large language model. Decoding, evaluated as word-grouped top-10 retrieval against permutation baselines, was reliably above chance, extended to mid-frequency and rare words, and scaled log-linearly with training-data volume with no sign of saturation. Removing occipital and posterior-temporal electrodes reduced the word-level gain by roughly one third but left context tracking unchanged. Control analyses separate word-level decoding from narrative context tracking and from a non-neural positional prior introduced by the transformer's positional embedding. These results establish that open-vocabulary word-level information is recoverable from EEG during silent reading, and that decoding is data-limited rather than saturated.

I. Marquardt, A. Alchanat, Priyanka Jain · 0 citations
Preprint Jul 2026

DS-MTNet:Structured Multi-Task EEG Decoding for Human-Machine Collaboration

Current human-machine collaboration (HMC) systems rely on environment-facing sensors to observe visible actions and scene states, but the internal perceptual, intention-related, and state-related processes of operators remain insufficiently integrated into machine perception. Electroencephalography (EEG) provides a non-invasive, time-resolved modality to capture neural activity associated with these processes and can serve as an additional sensing channel in HMC. However, HMC-relevant EEG evidence is often mixed in continuous recordings. Existing EEG decoding methods usually target task-specific classification or aggregate prediction, so multiple HMC-relevant readouts are rarely organized in a unified EEG representation. To address this gap, this paper proposed the Decomposed-Source Multi-Task Network (DS-MTNet), a structured multi-task EEG decoding framework. DS-MTNet integrated three streams, namely EEG waveforms, task-routed source embeddings, and temporal-spectral power features, into reusable slots and used dual gating mechanisms to route task-specific components. The model was tested on a sustained-attention driving EEG dataset with three representative readouts: lane-departure-related epochs for environmental-event processing, steering-response stage for response preparation, and reaction-time-defined alertness state for internal state. DS-MTNet achieved the best mean performance among traditional, single-task deep, and multi-task EEG baselines, with the most robust gains observed for steering-response stage decoding. Ablation and interpretability analyses suggested that DS-MTNet jointly decoded multiple readouts and organized event-related, response-related, and state-related EEG evidence in a unified source-slot representation. These findings provide a computational step toward incorporating operator-related neural evidence into machine perception in HMC.

Xinjia Yu, Yang Zhou, Jing Yang et al. · 0 citations