EEG-Based Schizophrenia Detection Under Subject-Wise Cross-Validation: A Multi-Method XAI Benchmark
Abstract
This study evaluates whether subject-wise cross-validation and multi-method explainability analysis can yield reproducible and neurophysiologically interpretable deep learning models for EEGbased schizophrenia detection. Schizophrenia is a severe psychiatric disorder for which no objective neurophysiological biomarker is yet established for routine clinical use. Electroencephalography (EEG) offers a cost-effective and temporally precise window into cortical dynamics, yet deep learning models for EEG-based schizophrenia classification are frequently evaluated with segment-level cross-validation that allows participant-specific patterns to leak into test sets, yielding inflated accuracy estimates. This study presents a comprehensive framework combining three architecturally distinct deep learning models — EEGNet, a CNN-LSTM hybrid, and a patch-based EEG Transformer — with six explainable artificial intelligence (XAI) methods spanning gradient-based (Saliency, Integrated Gradients, DeepLIFT), game-theoretic (SHAP), activation-based (Grad-CAM), and perturbation-based (LIME) paradigms. All models are evaluated under strict subject-wise 5-fold cross-validation on the publicly available 84-participant resting-state EEG dataset from M.V. Lomonosov Moscow State University, supplemented by a multi-component regularization pipeline. The CNN-LSTM model achieved the highest classification accuracy of 0.8192±0.0543 and AUC-ROC of 0.8776±0.0906. Crucially, XAI analyses were conducted across all five cross-validation folds, yielding 6×3×5 = 90 attribution analyses (6 XAI methods × 3 architectures × 5 folds, corresponding to 15 trained model instances each evaluated under 6 XAI methods). This multi-fold XAI design revealed that the right posterior temporal electrode T6 was identified as the most discriminative channel at the grand-average level in three of the five cross-validation folds, with a grand-average importance score of 0.977. All three model architectures and five of the six XAI methods independently converged on T6 as the top-ranked channel across 5-fold averages; LIME, while identifying T6 among the most important channels, showed greater variability with F8 (right frontal) as its top-ranked channel, consistent with its perturbation-based nature. This aggregate-level convergence, though weaker at the level of individual folds and models (see Section 4.5), provides evidence for the discriminative role of the right posterior temporal scalp region — overlying the superior temporal gyrus, a region critically implicated in auditory processing and auditory verbal hallucinations in schizophrenia, though scalp EEG cannot precisely localize the underlying cortical source.