Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best, and transfers across corrector families and adds zero parameters to the inference path.
Chao-Han Huck Yang, Zih-Ching Chen, Piotr Żelasko et al.· arXiv.org· 0 citations
This work proposes an efficient method to add streaming ASR capabilities to an existing duplex S2S model by introducing a lightweight ASR head in parallel to the agent text head, enabling real-time user transcription while preserving full-duplex conversational capabilities including turn-taking and barge-in handling.
Ke Hu, Nourchene Ferchichi, Edresson Casanova et al.· 1 citation
VoiceChat-TTS is proposed, a low-latency, continuous, and streamable text-to-speech model for interactive agents that enables always-on, responsive speech generation while preserving modularity and high speech quality.
Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor et al.· 3 citations
This work proposes an always-on attention-coordination layer that mediates this interface and allocates human attention across one or more working agents, and introduces JarvisBench to evaluate both directions of this coordination.
JarvisBench, a benchmark for measuring the dual value of mediation in long-horizon agent workflows, is introduced and preliminary results suggest that Jarvis-style mediation can provide trace-grounded responses to user questions and improve task performance when sparse user guidance is injected at appropriate moments.
Chen Chen, Zhehuai Chen· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.