Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best, and transfers across corrector families and adds zero parameters to the inference path.
Chao-Han Huck Yang, Zih-Ching Chen, Piotr Żelasko et al.· arXiv.org· 0 citations
This work proposes an efficient method to add streaming ASR capabilities to an existing duplex S2S model by introducing a lightweight ASR head in parallel to the agent text head, enabling real-time user transcription while preserving full-duplex conversational capabilities including turn-taking and barge-in handling.
Ke Hu, Nourchene Ferchichi, Edresson Casanova et al.· 1 citation
VoiceChat-TTS is proposed, a low-latency, continuous, and streamable text-to-speech model for interactive agents that enables always-on, responsive speech generation while preserving modularity and high speech quality.
Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.