SphereVAE: Hyperspherical Latent Autoencoders for Robust Autoregressive Speech Representation Modeling
With the rapid development of speech generation technology, discrete codec representations have been widely used because they provide a stable prediction paradigm. In expressive speech generation, however, the quantization bottleneck of discrete codecs results in information gaps in fine-grained prosody, timbre, pronun...