With the rapid development of speech generation technology, discrete codec representations have been widely used because they provide a stable prediction paradigm. In expressive speech generation, however, the quantization bottleneck of discrete codecs results in information gaps in fine-grained prosody, timbre, pronun...
Hao-Yu Zhang, Jing-Bin Hu, Han-Ke Xie et al.· 0 citations
Controllable synthesis of nonverbal vocalizations (NVVs) is es- sential for natural and expressive speech, but remains challeng- ing due to their acoustic diversity and imbalanced distribution in existing corpora. To address these challenges, we develop an NVV-aware DiTAR system that models continuous speech latents, e...
Zi-Yu Zhang, Yun Chen, Tai-Hui Wang et al.· 0 citations
Continuous-latent Autoregressive Diffusion Transformer (AR-DiT) models have demonstrated immense potential in zero-shot speech generation. However, they still suffer from limited decoding stability when synthesizing long utterances or complex linguistic structures. This instability primarily stems from a restricted his...
Zi-Yu Zhang, Tian-Lun Zuo, Han-Zhao Li et al.· 0 citations
It is shown that structured LLM prompting and scalar supervision collapse rubric dimensions, yielding near-zero correlation with human ratings and strong cross-dimension coupling, are effective for segment-level SI evaluation.
Zi-Yu Zhang, Satoshi Nakamura· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.