Skip to content
Preprint

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

Aug 2026 · 3 citations · 23 references
Engineering Computer Science

TL;DR

VoiceChat-TTS is proposed, a low-latency, continuous, and streamable text-to-speech model for interactive agents that enables always-on, responsive speech generation while preserving modularity and high speech quality.

Abstract

Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing multi-stage pipelines, but often compromise speech quality because accurate ASR, interruption handling, and high-fidelity synthesis must be optimized jointly. We propose VoiceChat-TTS, a low-latency, continuous, and streamable text-to-speech model for interactive agents. VoiceChat-TTS is driven directly by LLM text-token streams, supports explicit interruption via control tokens, and produces silence when no textual input is available. The model enables always-on, responsive speech generation while preserving modularity and high speech quality, and it supports mid-utterance interruptions without resetting the KV cache.

View source

Similar papers

#natural language process... Preprint Sep 2026

KABURI-TTS: Phoneme-Keyed Activity-conditioned Bi-channel Utterance Rendering for Interaction

Realizing full-duplex spoken dialogue requires large amounts of two-channel, one-speaker-per-channel conversational speech data. Although conversational text-to-speech (TTS) engines have been developed, they are not necessarily robust to two-party simultaneous phenomena such as backchannels, interruptions, and overlaps...

Ryuichiro Higashinaka, Shinnosuke Takamichi, Tetsuji Ogawa · 0 citations
Preprint Sep 2026

CharDuplex: Building Character-Consistent Full-Duplex Spoken Dialogue Models

Full-duplex speech models are moving voice interaction beyond conventional turn-taking, yet natural conversation is shaped not only by when an agent speaks, but also by how it behaves as a conversational character. We present CharDuplex, a character-driven full-duplex speech model that combines real-time spoken interac...

Donghang Wu, Yi-Si Liu, Chen Chen et al. · 0 citations
Preprint Sep 2026

RePlay: Retrieval-Based Voice Playback for Multi-Turn spoken dialogue

Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorded performance, while cascaded systems can be constrained to...

Sathvik Udupa, Naveen Kumar, Ryan Folmsbee · 0 citations
Conference Open access Sep 2026

DeepL Voice: Real-Time Speech-to-Speech Translation

DeepL Voice is a real-time speech-to-speech translation system for global business communication, following a pragmatic incremental approach: developing a production-grade cascaded speech-to-speech-translation (S2ST) system, while exploring end-to-end solutions in parallel. The production system (launched November 2024...

Johannes Ernesti, Peter Kaiser, Jonas Heinze et al. · 0 citations
#natural language process... Preprint Sep 2026

Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models

This work proposes an efficient method to add streaming ASR capabilities to an existing duplex S2S model by introducing a lightweight ASR head in parallel to the agent text head, enabling real-time user transcription while preserving full-duplex conversational capabilities including turn-taking and barge-in handling.

Ke Hu, Nourchene Ferchichi, Edresson Casanova et al. · 1 citation
Preprint Aug 2026

PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advance dialogue state faster than clients can play it, subsequent user speech may be interpreted based on content the user never heard. We call this failure Generative Context Mis-anchor...

Shi-Bo Wang, Zicheng Zhang, Libo Wang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.