Skip to content
Preprint

Brain2Speech-Net: Fast and Intelligible Brain-to-Speech Synthesis Without Text Decoding

Sep 2026 · 0 citations · 32 references
Engineering Computer Science

Abstract

The loss of speech limits communication for individuals with paralysis. Direct neural-to-speech synthesis is challenging due to the limited availability of neural data for training speech brain-computer interfaces. Most existing systems rely on cascaded neural-to-text-to-speech pipelines, which increase inference latency and propagate errors across stages. We present Brain2Speech-Net, a single-stage neural-to-speech generation framework without intermediate text decoding. We use a differentiable phoneme bottleneck and a deep-HMM alignment mechanism to map long neural recordings into the latent space of a text-to-speech (TTS) model, enabling high-quality speech synthesis. Brain2Speech-Net is the only system in our comparison that produces intelligible speech while generating faster than real time.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.