Discrete audio tokenization has become the critical interface between raw waveforms and autoregressive modeling in recent music generation. As a result, music tokenizers must simultaneously support high-fidelity reconstruction and produce discrete sequences that remain amenable to language modeling. Existing reconstruc...
Hua-Kang Chen, Guo-Bin Ma, Yue-Peng Jiang et al.· 0 citations
Dynamic speaking style adaptation in multi-turn multimodal interaction remains a major challenge for text-to-speech (TTS) systems. Existing context-aware TTS (CTTS) methods typically map dialogue context to speech in an end-to-end manner. Such implicit modeling makes contextual style decisions difficult to supervise, w...
Wen-Jie Tian, Kang-Xiang Xia, Jing-Bin Hu et al.· 0 citations
With the rapid development of speech generation technology, discrete codec representations have been widely used because they provide a stable prediction paradigm. In expressive speech generation, however, the quantization bottleneck of discrete codecs results in information gaps in fine-grained prosody, timbre, pronun...
Hao-Yu Zhang, Jing-Bin Hu, Han-Ke Xie et al.· 0 citations
Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-f...
Kang-Xiang Xia, Xin-Fa Zhu, Hang-Rui Hu et al.· 0 citations
Recent advancements in discrete token-based speech generation have highlighted the importance of efficient token-to-waveform synthesis in streaming and dialogue scenarios. Flow-matching acoustic decoders achieve high-quality token-to-mel generation, but their iterative sampling requires multiple neural function evaluat...
Han-Ke Xie, Xia-Ming Ren, Qi-Rui Zhan et al.· 0 citations
The ISCSLP 2026 CoT-TTS Challenge requires TTS systems to generate Chain-of-Thought (CoT) reasoning from dialogue history before synthesizing contextually appropriate speech. While the official baseline establishes a unified architecture, it remains constrained by limited contextual comprehension, weak instruction fide...
Jing-Bin Hu, Lu-Yu Wang, Wen-Jie Tian et al.· 0 citations
SemBridge is proposed, a training-only semantic-token anchoring framework for continuous-latent autoregressive speech generation and demonstrates that explicit semantic-token supervision for autoregressive state learning is an effective and general direction for continuous speech generation.
Han-Ke Xie, Haopeng Lin, J. Qian et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.