English-only benchmarks expose only a narrow slice of voice-agent behavior. We introduce $\tau$-Multilingual, extending $\tau$-Voice to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin with native-speaker review and evaluation of generated language and spoken output. Across 4,500 full-duplex calls and five vo...
Soham Ray, Edgard dos Santos Paiva, Rubén Valenzuela et al.· 0 citations
A 200-task voice benchmark spanning 10 entity types, controlled difficulty, caller realisms, and three environments is introduced, identifying strategy selection and successful recovery as the central bottlenecks in exact spoken entity collection.
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a...
Shi-Xiu Quan, Keshav Dhandhania, Karthik R. Narasimhan et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.