TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context
This work presents TontaubeV1, a model that preserves natural prosody while enabling streaming from a single consumer GPU, and is designed primarily for English and German, with additional multilingual support.
Fritz Cremer, Jonathan Cremer
· 0 citations