The loss of speech limits communication for individuals with paralysis. Direct neural-to-speech synthesis is challenging due to the limited availability of neural data for training speech brain-computer interfaces. Most existing systems rely on cascaded neural-to-text-to-speech pipelines, which increase inference laten...
Brain2Speech-Net is presented, among the first single-stage frameworks to remain intelligible under limited data while removing intermediate text decoding, and achieves strong intelligibility in objective and listening tests while running faster than real time.
DiffAnon is proposed, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation, and is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control.
Ismail Rasim Ulgen, Zexin Cai, Nicholas Andrews et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.