Skip to content

ConversationalVoice: Full-Duplex Speech Data from Real Conversations through Source-Faithful Reconstruction and Conversation-Grounded Expansion

Sep 2026 · 1 citation · 12 references
Computer Science Engineering

TL;DR

Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts, is presented, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts.

Abstract

Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordings. We present Conversational Voice, a pipeline that converts real two-speaker excerpts into three complementary training-data artifacts. (1) Separation recovers speaker-specific tracks with stable speaker assignments, a canonical transcript, and naturally observed interaction timing. (2) Reconstruction generates speech in matched voices from a fixed source transcript, reconstructs the source turn order, pauses, and overlaps, and adds word-level alignment and delivery instructions. (3) Expansion generates new dialogue constrained by the source context, speakers, and observed interaction pattern. Automatic speaker-verification metrics remain strong across stages, with same-speaker similarity of 0.983-0.991 and positive discrimination margins of 0.199-0.209. Predicted speech quality (NISQA MOS) is 3.56 for separation, 4.41 for reconstruction, and 4.61 for expansion. A Gemini-based automatic evaluator assigns expansion mean scores of 4.94/5 for contextual coherence and 4.80/5 for dialogue naturalness. Expansion and reconstruction exhibit broadly similar interaction profiles; expansion's turn, overlap-event, backchannel, and interruption rates are 4.6%, 8.0%, 13.2%, and 16.0% lower, respectively. We evaluate data properties only; downstream gains in full-duplex model training remain for future work.

View source

Similar papers

An automated pipeline for standardised speech-unit annotation in spontaneous dialogue

Quantifying conversational dynamics requires reliable identification of interactional units and their temporal boundaries, but speech activity alone does not distinguish conversational turns from listener feedback or within-turn pauses. We present an automated pipeline for extracting turns and backchannels from separat...

Han-Lu He, H. V. Skat-Rørdam, I. Örnólfsson et al. · 0 citations

DuplexGen: Decoupling Content, Timing, and Acoustics for Synthetic Dialogue Speech

D DuplexGen is presented, a dialogue synthesis framework that explicitly decouples content, timing, and acoustics and produces conversational dynamics closer to the real-dialogue reference distribution than the stitching baselines evaluated in this study.

Pengcheng Wang, Sheng Li, JiyiLi et al. · 0 citations
#natural language process... Preprint Sep 2026

DuplexDrama: A Synthesized Dialogue Dataset with Scenarios, Full-Duplex Behaviors, Expressive Speech, and Sound Events

We present DuplexDrama, the first synthesized spoken dialogue dataset that simultaneously covers four dimensions: (i) complete persona and scenario settings; (ii) three full-duplex behaviors (interruption, backchannel, incomplete); (iii) expressive speech with persona-aligned emotion labels; and (iv) script-aware sound...

Qing-Xiang Guo, Wen-Ke Fan, Shuo-Feng Zhao et al. · 2 citations
#natural language process... Preprint Sep 2026

KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction

Realizing full-duplex spoken dialogue requires large amounts of two-channel, one-speaker-per-channel conversational speech data. Although conversational text-to-speech (TTS) engines have been developed, they are not necessarily robust to two-party simultaneous phenomena such as backchannels, interruptions, and overlaps...

Ryuichiro Higashinaka, Shinnosuke Takamichi, Tetsuji Ogawa · 0 citations
#artificial intelligence Preprint Sep 2026

A Harness for Synthesizing Diverse Naturalistic Full-Duplex Conversations

Full-duplex dialogue systems, which listen while speaking, must distinguish a completed turn from a pause within a turn and an interruption that requests a turn from a brief acknowledgment or speech addressed to a third party. Yet existing conversational corpora provide limited control over these events and limited lab...

Matthew Z. Sun, Vinay Kothapally, Meng Yu et al. · 0 citations
Preprint Sep 2026

Dialogue-Based Streaming Audio-Visual Target Speaker Extraction with Predictive Dialogue Information

In face-to-face, real-time communication, a talking agent must track the target speaker through natural pauses, turn-taking, and backchannels, often amid background cross-talk. Most target speaker extraction (TSE) studies, however, rely on simulated mixtures with full or sparse overlap and ignore the turn-taking of rea...

Shu-Han Zhang, Wen-Xuan Wu, Hai-Zhou Li · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.