Efficient fMRI and Textual Alignment for Image Reconstruction from Human Brain Activity
Abstract
In neural decoding research, reconstructing natural images from fMRI signals poses a captivating yet challenging problem. Conventional approaches use basic linear mapping functions to project fMRI signals into a prior latent space (e.g., image and text embeddings) and subsequently utilize a pre-trained image generation model to create images conditioned on the embeddings. While effective at capturing low-level features like layout, texture, and shape, these methods often fail to capture high-level features such as object categories, spatial positions, and the number of objects. These failures are often due to the insufficient semantic information carried by the text embeddings using simple mapping functions. In this work, we, therefore, concentrate on improving text embeddings to address the above challenges. Specifically, we propose the MDIR framework, which employs a unified and deep neural network called MindMapper to align fMRI signals with text embeddings more effectively. As a result, our approach enhances the consistency between reconstructed images and the semantic information in text captions. This method achieves superior semantic fidelity, producing images closely resembling ground truth in fine-grained detail. Experiments on standard fMRI-Reconstruction benchmark show significant improvements compared to baseline method in both qualitative and quantitative results.