D DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding, and shows that even the best voice agents leave substantial room for improvement on all 3 axes.
Abstract
Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation. To tackle this DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.
Multiparty Bench is introduced, the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts and assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness.
Yi-Jen Shih, S. Kuan, Guan-Ting Lin et al.· 2 citations· ⚡1
Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow and per-call cost makes on-policy reinforcement learning prohibitive, or stay in te...
Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or ta...
T. Chiba, Guang-Zhi Sun, Zhe-Qi Yuan et al.· 0 citations
The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable performance as general voice assistants, their performance in specialized domains remains underexplore...
He-Yang Liu, Jia-Yi Huang, Wen Xiao et al.· 0 citations
Hear2Act is introduced, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes that show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do...
Xin-Yi Liu, H. Nayyeri, Dilek Hakkani-Tur et al.· 3 citations· ⚡1
Conversation Coach is proposed, a voice-first AI system that enables managers to rehearse difficult workplace conversations in a realistic spoken format and offers superior reasoning essential for coaching quality.
Wu Fanyou, Suraj Maharjan, Ainur Yessenalina et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.