Skip to content

MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant

Sep 2026 · 2 citations · ⚡ 1 influential · 82 references
Engineering Computer Science

TL;DR

Multiparty Bench is introduced, the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts and assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness.

Abstract

Conversational voice agents have advanced significantly, offering increasingly natural human-machine interactions through both cascaded and end-to-end architectures. However, while recent benchmarks extensively evaluate dyadic interactions and passive audio comprehension, they largely overlook a prevalent real-world scenario: multi-party conversations. Evaluating agents in these settings is fundamentally more challenging than in dyadic interactions due to the exponentially greater conversational complexity. For voice agents to integrate seamlessly into human group dynamics, they must not only generate contextually appropriate responses but also demonstrate a nuanced understanding of open turn-taking. To address this gap, we introduce Multiparty Bench (MP-Bench), the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts. MP-Bench assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness. Additionally, we incorporate comprehension-based question-answering tasks as a complementary evaluation. By benchmarking 12 voice agents, we find that real-time voice agents stay at or below 22% on multiparty comprehension and remain near chance on multiparty turn-taking, exposing an open challenge for real-time voice agents under multiparty scenario.

View source

Similar papers

#natural language process... Preprint Sep 2026

MSI-Bench: Evaluating Multi-Speaker Voice Interaction for Collaborative AI Agents

Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, and collaborative work, are inherently multi-speaker. Supporting these settings introduces challenges that are largely absent from one-on-one interaction. We introduce the...

Chenxu Xiong, Dong-Ming Shen, Yu-Zhi Tang et al. · 3 citations · ⚡1
Preprint Aug 2026

DuplexWorld: Can voice agents help you get through the day?

D DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding, and shows that even the best voice agents leave substantial room for improvement on all 3 axes.

Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli et al. · 2 citations
Preprint Sep 2026

Duplex-MPE: Benchmarking Multi-Party Interaction in Full-Duplex Dialogue

Real-time full-duplex speech models can listen while speaking, enabling natural interaction without rigid turn boundaries. Existing benchmarks evaluate turn-taking, interruption handling and multi-round dialogue, but largely centre on a designated user rather than an assistant participating in a shared conversation amo...

Chengqian Ma, Wen-Hao Feng, Wei-Xuan Jin et al. · 0 citations
#natural language process... Preprint Aug 2026

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

TurnBench, a multi-domain benchmark that pairs a 30-hour, hand-labeled corpus of dyadic human conversation with a standardized evaluation protocol for end-of-turn and interruption detection, finds end-of-turn recall stable across types, while interruption false positives are strongly type-dependent and concentrated in...

Freeman Jiang, Ramon Sanabria, Soham Deshmukh et al. · 5 citations · ⚡1

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents, and concludes with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cu...

S. Sara, Zien Sheikh Ali, Hunzalah Hassan Bhatti et al. · 3 citations
#artificial intelligence Preprint Sep 2026

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

Speech-to-speech dialogue models increasingly support persona control, yet existing spoken role-playing benchmarks remain largely character-centric and short-horizon. This leaves open whether spoken dialogue models can sustain diverse roles over extended interactions, especially beyond predefined fictional characters....

Yu-Qi Wang, Feng-Yuan Liu, Hao-Chen Luo et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.