D DuplexGen is presented, a dialogue synthesis framework that explicitly decouples content, timing, and acoustics and produces conversational dynamics closer to the real-dialogue reference distribution than the stitching baselines evaluated in this study.
Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dialogue synthesis pipelines typically generate dialogue content first and then insert interruptions, overlap, and backchannels using handcrafted markers or timing rules, makin...
Pengcheng Wang, Sheng Li, Ji-Yi Li et al.· 0 citations
Full-duplex dialogue systems, which listen while speaking, must distinguish a completed turn from a pause within a turn and an interruption that requests a turn from a brief acknowledgment or speech addressed to a third party. Yet existing conversational corpora provide limited control over these events and limited lab...
Matthew Z. Sun, Vinay Kothapally, Meng Yu et al.· 0 citations
We present DuplexDrama, the first synthesized spoken dialogue dataset that simultaneously covers four dimensions: (i) complete persona and scenario settings; (ii) three full-duplex behaviors (interruption, backchannel, incomplete); (iii) expressive speech with persona-aligned emotion labels; and (iv) script-aware sound...
Qing-Xiang Guo, Wen-Ke Fan, Shuo-Feng Zhao et al.· 2 citations
Full-duplex speech models are moving voice interaction beyond conventional turn-taking, yet natural conversation is shaped not only by when an agent speaks, but also by how it behaves as a conversational character. We present CharDuplex, a character-driven full-duplex speech model that combines real-time spoken interac...
Donghang Wu, Yi-Si Liu, Chen Chen et al.· 0 citations
Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts, is presented, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts.
Richard Yucheng He, Bao-Dong Cao, Chen Xu et al.· 1 citation
A controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement, and unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration.
Kang-Xiang Xia, Xin-Fa Zhu, Hang-Rui Hu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.