Skip to content

Author

Zhehuai Chen

We have 6 of 46 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Voice Memory for Agentic Speech Recognition

Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best, and transfers across corrector families and adds zero parameters to the inference path.

Chao-Han Huck Yang, Zih-Ching Chen, Piotr Żelasko et al. · 0 citations
#natural language process... Preprint Sep 2026

Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models

This work proposes an efficient method to add streaming ASR capabilities to an existing duplex S2S model by introducing a lightweight ASR head in parallel to the agent text head, enabling real-time user transcription while preserving full-duplex conversational capabilities including turn-taking and barge-in handling.

Ke Hu, Nourchene Ferchichi, Edresson Casanova et al. · 1 citation
Preprint Aug 2026

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

VoiceChat-TTS is proposed, a low-latency, continuous, and streamable text-to-speech model for interactive agents that enables always-on, responsive speech generation while preserving modularity and high speech quality.

Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor et al. · 3 citations
Preprint Aug 2026

JarvisBench: Always-on Intelligence Between Humans and Agents

This work proposes an always-on attention-coordination layer that mediates this interface and allocates human attention across one or more working agents, and introduces JarvisBench to evaluate both directions of this coordination.

Chen Chen, Zhehuai Chen · 0 citations
Jul 2026

Just A Rather Very Intelligent Spoken Agent

JarvisBench, a benchmark for measuring the dual value of mediation in long-horizon agent workflows, is introduced and preliminary results suggest that Jarvis-style mediation can provide trace-grounded responses to user questions and improve task performance when sparse user guidance is injected at appropriate moments.

Chen Chen, Zhehuai Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.