Skip to content

Cross-Cohort Spectral-Temporal Dissociation in Frozen EEG Foundation-Model Representations

Jul 2026 · 1 citation · 26 references
Biology Computer Science

Abstract

Objective. We tested whether frozen representations from five EEG foundation models support decoding of long-range temporal correlations, measured as the detrended-fluctuation-analysis (DFA) exponent of the alpha-band amplitude envelope. Approach. REVE, LaBraM, BENDR, CBraMod, and BIOT were evaluated in CAUEEG and BrainLat. A common 240 s estimator used 8-13 Hz filtering, DFA over 2-23.8 s, artifact masking, and quality control. One fixed nested-cross-validation readout predicted DFA and a fixed-mode aperiodic exponent. Controls tested pre-pool order sensitivity and aperiodic residualization. Results. CAUEEG included 764 recordings and BrainLat 79. BIOT decoded DFA in CAUEEG (R-squared = 0.232; conditional subject-bootstrap 95 percent interval, 0.121-0.310), and CBraMod was positive but imprecise (R-squared = 0.121; 0.003-0.214). Neither replicated in BrainLat, where all five point estimates were negative. In contrast, CBraMod and BIOT decoded the aperiodic exponent in both cohorts (R-squared = 0.459-0.757). BIOT remained positive after removal of the measured linear aperiodic association in matched CAUEEG data (R-squared = 0.240). The post-hoc order control was batch- and configuration-sensitive. Because chronological EEG epochs are not exchangeable, it was descriptive, not an LRTC-specific test. No revised DFA transfer direction passed source-label permutation testing. Cohort membership was near-ceiling decodable from all five embeddings, but this is not a pure site effect. Significance. CBraMod and BIOT show a replicated, model-specific spectral-temporal dissociation: aperiodic decoding is present in both cohorts, whereas alpha-envelope DFA decoding is cohort-dependent. These findings bound the evaluated readouts; they do not establish representational absence or an architectural cause. Transfer and clinical associations remain exploratory.

View source

Similar papers

Preprint Jul 2026

Bayesian Complete-Pooling in Cross-Subject Classification for Motor Imagery Electroencephalogram

Brain-computer interfaces (BCIs) have long sought calibration-free operation, but classifiers are typically benchmarked by discrimination alone, blind to whether predicted probabilities are well calibrated - a meaningful gap given nonstationary electroencephalogram (EEG) signals and the risk of overconfident point-estimate classifiers under distribution shift. We conducted a large-scale study contrasting Bayesian complete-pooling models against frequentist baselines for cross-subject, left-hand versus right-hand motor imagery EEG classification across 20 datasets. Six frequentist pipelines were each paired with an analogous Bayesian pipeline sharing identical feature engineering, fit via Markov chain Monte Carlo posterior sampling. Our primary metric was the Brier score, decomposed into reliability and resolution, alongside AUROC for discrimination and Shannon entropy for sharpness. Each metric was analyzed via random-effects meta-analysis (REML, Knapp-Hartung adjustment), verified by leave-one-out influence analysis. Bayesian complete-pooling produced statistically but not practically significant improvements in reliability and increases in predictive uncertainty (lower sharpness); Brier score, resolution, and discrimination showed no significant differences. Between-study heterogeneity was low across all metrics, though the reliability result was sensitive to leave-one-out removal. We additionally profiled computational cost, finding that Bayesian pipelines consumed roughly thirteen times more energy than their frequentist counterparts, a cost that remains modest relative to common household appliances. These results suggest that Bayesian complete-pooling alone offers limited practical benefit for cross-subject motor imagery classification, and that partial-pooling across subjects and sessions is a more promising direction for future work.

Ethan Davis · 0 citations
Open access Aug 2026

Reliability and disease sensitivity are dissociable properties of EEG foundation-model representations

EEG foundation models (EEG-FMs) are evaluated almost entirely on disease-discrimination accuracy. A clinical biomarker additionally requires measurement reliability, the stability of repeated measurements on the same individual, which regulatory biomarker frameworks treat as a prerequisite that discrimination does not imply. We asked whether frozen EEG-FM representations provide such stability, whether it is predictable from conventional model descriptors, and what information supports it. We measured test-retest reliability, disease discrimination, and representation distinctiveness for nine frozen representations: six EEG-oriented foundation models spanning masked, contrastive and predictive pretraining, handcrafted spectral features, and two general-purpose time-series models with no EEG exposure. All were evaluated under one preprocessing pipeline across two healthy retest cohorts, at roughly one month and two years, and three neurodegenerative cohorts. Reliability was measured in healthy adults only; the disease cohorts contribute cross-sectional discrimination. Reliability, quantified as the intraclass correlation coefficient (ICC), varied enormously (mean 0.08 to 0.76; coeffiicient of variation, CV, 53.0%) while disease discrimination, measured as area under the receiver operating characteristic curve (AUC), occupied a far narrower range across the same nine (AUC CV 5.4%), a roughly tenfold difference in relative dispersion, described rather than formally tested. We did not test formal AUC equivalence, so we describe discrimination as varying substantially less than reliability rather than as equivalent. The variation was not consistently explained by pretraining paradigm or domain among the models studied, and a model with no EEG exposure was among the most reliable tested. Alpha-band information contributed disproportionately to reliability, whereas theta-band information ranked first for Alzheimer’s disease and frontotemporal dementia discrimination, directionally consistent with established EEG evidence in both conditions. Only the reliability half of that contrast is individually significant. Subspace geometry and band ablation, two methodologically distinct analyses, both indicate that the two properties are partially, not fully, dissociable, and network architecture determines whether the dissociation is preserved, traded off, or jointly degraded across depth. Reliability showed no detectable association with discrimination, pretraining paradigm, or domain, and had to be measured directly. We recommend it become a standard evaluation axis for EEG-FM representations intended for longitudinal or biomarker use, and release a reproducible pipeline.

B. Gebregergis, Haben Yhdego, Tewolde Teklu · 0 citations
Conference Jul 2026

Ocular Artifact Removal in EEG: A Review

Electroencephalography (EEG) offers millisecond-scale temporal resolution for studying brain activity, yet its practical value is persistently compromised by ocular artifacts—unwanted electrical potentials from eye blinks and saccadic movements that can exceed genuine cortical signals by an order of magnitude. Because these artifacts share spectral content with neurologically meaningful oscillations, simple frequency-domain filtering proves inadequate. This survey covers four decades of suppression research: from regression-based subtraction and matrix factorisation (PCA, ICA) through signal-adaptive transforms (wavelet shrinkage, EMD variants) to contemporary deep-learning architectures including CNNs, GANs, and transformers. Work published from 2021 to 2026 receives particular emphasis, covering foundation-model pre-training, self-supervised artifact rejection, and edge-deployable pipelines for real-time brain–computer interface (BCI) applications. Performance is benchmarked via ∆SNR, %BAR, and RRMSE across five comparative tables and two trend figures. Ongoing challenges—ground-truth scarcity, inter-subject variability, edge-deployment constraints, and the absence of standardised benchmarks—are examined alongside prospective research directions.

A. Mishra, M. Choudhry · 0 citations
Conference Jul 2026

Individual Calibration for Mental Fatigue-Related State Recognition Using Five Frontal Dry-Electrode EEG Channels

This study investigated whether five frontal dry-electrode electroencephalography (EEG) channels can support recognition of mental fatigue-related states after individual calibration. Thirteen healthy adult men completed pre- and post-rest recordings, N-back, Stroop, and a 15-min sustained attention to response task (SART). EEG was recorded from FP1, FP2, AF7, FPz, and AF8. Following quality control, 12 participants were retained for EEG analyses. Spectral, Hjorth, entropy, asymmetry, and inter-channel correlation features were extracted from the early and late SART periods. A subject-dependent nested contiguous temporal-block cross-validation framework selected temporal aggregation, feature set, classifier, and feature number using training data only. KSS increased from 1.77±0.73 to 5.15±0.55 after SART (p < 0.001). No-go commission errors increased from 25.44%±16.63% to 35.87% ± 20.85% (p = 0.040), whereas Go median reaction time decreased (p < 0.001). The main model achieved a mean balanced accuracy (BA) of 0.831±0.124, a median BA of 0.850, and a mean area under the receiver-operating-characteristic curve (AUC) of 0.920±0.114. Chronological holdout testing yielded a mean BA of 0.774±0.204, whereas strict cross-subject classification remained near chance. These findings support the use of five frontal dry-electrode EEG channels for subject-dependent monitoring of mental fatigue-related states after individual calibration, but do not yet support general classification without target-subject calibration.

Chao Chen, Xianjin Shi, Dongyue Wu · 0 citations
Open access Jul 2026

Back to the Future of qEEG: Lifespan Normative Modeling of Spectral Ratios and Functional Indices with Potential Applications to Therapeutic Monitoring

Quantitative EEG (qEEG) provides objective, millisecond-resolution measures of brain dynamics. Despite decades of methodological advances, clinically relevant derived indices—spectral power ratios, cognitive-emotional state markers, and physiological parameters—are typically reported as raw values without the normative context required for individualized clinical inference. To develop the first systematic age-dependent normative models for this family of derived qEEG indices using a multinational database, enabling probabilistic Z-score interpretation at the individual level with potential applications in objective therapeutic monitoring. Normative modeling was applied to the HarMNqEEG database (n = 1,564 neurologically healthy participants, ages 5–97, 9 countries, eyes-closed resting state). Electrode-level Spectral Normalization (ESN) removed inter-individual and inter-device amplitude variability while preserving the neurophysiological interpretability of each index. Age-dependent normative trajectories were estimated using Generalized Additive Models for Location, Scale, and Shape (GAMLSS) with P-splines on log(age), allowing conditional mean and variance to vary non-linearly across the lifespan. GAMLSS modeling revealed significant non-linear age-dependent trajectories for all indices. Slow-wave-dominated ratios showed steep decreases from childhood to early adulthood, consistent with cortical maturation; alpha-dominated indices increased during adolescence before stabilizing. ESN normalization yielded well-calibrated normative residuals across the full age range for all indices, except Valence. For the Valence index, a quantile regression model is provided as the recommended normative reference due to its spike-and-slab marginal distribution. These normative models provide a principled, age-adjusted probabilistic framework for individual-level qEEG interpretation, based on eyes closed resting state recordings, laying the methodological groundwork for future clinical validation in diagnostic and therapeutic monitoring applications. The ESN strategy requires no knowledge of recording equipment, ensuring broad applicability across clinical and research settings.

J. Bosch-Bayard, J. Guerrero-Sauzameda, R. I. Bosch-Bayard et al. · 0 citations
Preprint Jul 2026

Subject-Level Heterogeneity in EEG Motor Imagery Decoding: A Large-Scale Benchmark and Portfolio-Based Reduction of the Search Space

Robust EEG motor imagery decoding remains limited by strong inter-individual variability, making it difficult to identify pipelines that generalize across users. We present a large-scale, standardized within-session benchmark of decoding pipelines across three public datasets: Cho2017 (52 subjects), PhysionetMI (109 subjects), and Zhou2016 (4 subjects). Using a common MOABB LeftRightImagery setting, two frequency bands (8-15 Hz and 8-30 Hz), and a broad combination of feature extraction, preprocessing, and classification steps, we analyzed 216,714 raw evaluation rows, which after structured aggregation yielded 44,928, 109,000, and 4,192 subject-level observations respectively. Covariance tangent-space projection (cov-tgsp) and Common Spatial Patterns (CSP) consistently defined the strongest methodological families, though their relative ordering was dataset-dependent. On Cho2017, the best family-level mean accuracy came from cov-tgsp in 8-30 Hz (0.712 +/- 0.140), whereas Zhou2016 favored CSP (0.832 +/- 0.121 in 8-15 Hz). These aggregate rankings concealed substantial subject-level heterogeneity: 42 distinct winning pipelines across 52 Cho2017 subjects, and 93 across 109 PhysionetMI subjects. We then used the benchmark as an empirical performance landscape for building compact portfolios of pipelines of size K. Several construction procedures were compared, including a ranking-based Top-K Mean heuristic and search-based strategies. Results were broadly consistent, with Top-K Mean giving the best trade-off. A single best global pipeline already retained 94.2% of the oracle in Cho2017 and 81.8% in PhysionetMI; at K = 12, oracle retention rose to 96.5% and 90.0%. The landscape is therefore subject-dependent, and this heterogeneity can be exploited through compact portfolios that make personalization more feasible.

Xavier Vasques, Paul Barbaste, Olivier Oullier · 0 citations

Related blog posts