English-only benchmarks expose only a narrow slice of voice-agent behavior. We introduce $\tau$-Multilingual, extending $\tau$-Voice to Spanish, Brazilian Portuguese, Hindi, Korean, and Mandarin with native-speaker review and evaluation of generated language and spoken output. Across 4,500 full-duplex calls and five vo...
Soham Ray, Edgard dos Santos Paiva, Rubén Valenzuela et al.· 0 citations
Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a dataset of 4,270 caller utterances from 250 phone calls to an...
A 200-task voice benchmark spanning 10 entity types, controlled difficulty, caller realisms, and three environments is introduced, identifying strategy selection and successful recovery as the central bottlenecks in exact spoken entity collection.
Soham Ray, Victor Barrès· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.