Skip to content
Preprint

DuplexWorld: Can voice agents help you get through the day?

Aug 2026 · 2 citations · 20 references
Computer Science

TL;DR

D DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding, and shows that even the best voice agents leave substantial room for improvement on all 3 axes.

Abstract

Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation. To tackle this DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

Multiparty Bench is introduced, the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts and assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness.

Yi-Jen Shih, S. Kuan, Guan-Ting Lin et al. · 2 citations · ⚡1
Preprint Aug 2026

SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning

Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing frameworks either cascade TTS and ASR around a proprietary voice API, where gradients cannot flow and per-call cost makes on-policy reinforcement learning prohibitive, or stay in te...

Jia-Jun Fan, Jing-Yuan Li, Prashanth Gurunath Shivakumar et al. · 1 citation · ⚡1
#artificial intelligence Preprint Sep 2026

Talk2Agent: Benchmarking Voice Interfaces for Text Agents

Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities, constraints, or ta...

T. Chiba, Guang-Zhi Sun, Zhe-Qi Yuan et al. · 0 citations
#natural language process... Preprint Sep 2026

$S^3$-Bench: Evaluating Speech Interaction Models as Scientific Voice Assistants

The advance of multimodal large language models (MLLMs) has fundamentally reshaped the paradigm of human-computer interaction, especially speech interaction models capable of seamless conversations. Despite remarkable performance as general voice assistants, their performance in specialized domains remains underexplore...

He-Yang Liu, Jia-Yi Huang, Wen Xiao et al. · 0 citations
Preprint Aug 2026

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

Hear2Act is introduced, a unified evaluation protocol for text and spoken assistants with 480 persona-grounded scenarios, hidden user concerns, and objectively verifiable outcomes that show that prosody matters when lexical evidence is insufficient, and that audio-capable LLMs can recover information from speech but do...

Xin-Yi Liu, H. Nayyeri, Dilek Hakkani-Tur et al. · 3 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.