Skip to content
Open access

Visual Stimulus Reconstruction from EEG Signals Using Masked Variational Autoencoders and Stable Diffusion

Aug 2026 · Journal of Intelligent Decision Making and Information Science · 0 citations · 18 references

Abstract

Electroencephalography (EEG)-based Brain–Computer Interfaces (BCIs) have emerged as a promising paradigm for decoding neural activity and translating brain signals into meaningful outputs. Among the various applications of BCIs, reconstructing or generating visual stimuli from neural signals represents a challenging and rapidly growing research direction. Recent advances in deep generative models, particularly diffusion-based image generation frameworks, have demonstrated remarkable capabilities in synthesizing realistic images from high-level semantic representations. Leveraging these developments, this work presents an EEG-guided image generation framework that combines EEG signal processing, latent representation learning, multimodal embedding alignment, and diffusion-based image synthesis. EEG data were collected from participants exposed to multiple categories of visual stimuli. The recorded signals were preprocessed using band-pass filtering, Artifact Subspace Reconstruction (ASR), Independent Component Analysis (ICA), and average re-referencing to remove physiological and environmental artifacts. A Masked Variational Autoencoder (VAE) with a Vision Transformer backbone was employed to learn robust latent EEG representations while preserving contextual neural information. To bridge the modality gap between EEG signals and visual representations, a CLIP-based embedding alignment strategy was introduced to project EEG features into a semantically meaningful image embedding space. These aligned EEG embeddings were subsequently used to condition a pretrained Stable Diffusion model through cross-attention mechanisms, enabling image synthesis directly from neural activity. Experimental results demonstrate that EEG signals contain sufficient information to recover coarse semantic characteristics of perceived visual stimuli. The proposed framework successfully generates images that preserve category-level visual information and exhibit meaningful correspondence with the observed stimuli. The study highlights the potential of integrating EEG-based neural decoding with modern generative AI models for future applications in assistive communication, cognitive state visualization, neurofeedback systems, and next-generation brain–computer interfaces.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.