Extensive experiments show that MSBraM achieves superior performance on other state-of-the-art pretrained models, demonstrating strong generalization and transferability, and indicate that explicitly modeling multi-scale temporal dynamics is critical for effective EEG foundation models.
Abstract
Self-supervised foundation models have recently shown strong potential for electroencephalogram (EEG)-based analysis. However, existing approaches struggle to capture the inherently multi-scale temporal structure of EEG signals, where local neural patterns and long-range dependencies jointly encode task-relevant information. This limitation hampers cross-scale representation learning and generalization across diverse downstream tasks. To address this challenge, we propose MSBraM, a Multi-Scale self-supervised Brain foundation Model designed to learn hierarchical EEG representations. MSBraM follows a two-stage pretraining framework. First, a multi-scale neural tokenizer discretizes raw EEG signals into semantic codes at different temporal resolutions via vector-quantized reconstruction. Second, the model is pretrained to predict masked codes using a curriculum multi-scale masking strategy, progressively integrating fine-grained local patterns with global temporal context. We pretrain MSBraM on over 2,400 hours of EEG data and evaluate it across 10 downstream tasks on 12 public datasets. Extensive experiments show that MSBraM achieves superior performance on other state-of-the-art pretrained models, demonstrating strong generalization and transferability. These results indicate that explicitly modeling multi-scale temporal dynamics is critical for effective EEG foundation models.
This study proposes the squeeze-and-excitation (SE)-data-adaptive Gaussian average filtering (DAGAF) adaptive tokenizer (SEDAT), a hybrid framework integrating SE-based spatial aggregation, DAGAF-based signal decomposition, instantaneous-frequency-guided adaptive segmentation, and Fourier-domain resampling into a singl...
Muhammad Zulkifal Aziz, Yue Zhuo, Bin-Wen Huang et al.· Journal of Neural Engineerin...· 0 citations
Graph convolutional networks (GCNs), owing to their capability to effectively process non-Euclidean data, have become one of the dominant approaches for decoding emotional states from electroencephalography (EEG) signals. However, current brain region division strategies exhibit limitations in information extraction, a...
Decoding neural activity into natural speech is a frontier in cognitive neuroscience and brain-computer interfaces. Existing methods overlook two critical issues when decoding continuous long-sequence speech: the pronounced inter-subject variability and substantial temporal fluctuations within individual EEG signals, a...
Cun-Hang Fan, Hui-Yao Lv, Sheng Zhang et al.· ACM Transactions on Autonomo...· 1 citation
Electroencephalography (EEG) is recorded continuously over hours, with relevant dynamics spanning timescales from milliseconds to hours. Most EEG foundation models nevertheless process fixed windows independently, limiting their ability to capture information encoded in long-timescale dynamics. State-space architecture...
Yi-Fan Wang, Hai-Ping Liu, Yang Cui et al.· 0 citations
Motor imagery–based brain–computer interface (MI-BCI) provides an effective pathway for post-stroke motor rehabilitation by decoding motor intentions from EEG signals. However, altered sensorimotor rhythms, nonlinear spatio-temporal dynamics, and limited clinical data make robust and generalizable post-stroke EEG decod...