Skip to content
Preprint

CoCoBench: A Cooperative Coordination Benchmark for Embodied Multi-Agent Task Planning

Aug 2026 · 0 citations · 31 references
Computer Science

TL;DR

CoBench is introduced, a construct-level benchmark for evaluating multi-agent embodied coordination in executable household tasks and shows that coordination ability is highly construct-specific: strong overall performance does not imply balanced competence across different coordination types.

Abstract

Agent systems powered by multimodal large language models (MLLMs) have advanced rapidly in recent years, yet existing embodied-agent benchmarks still lack fine-grained diagnostics for multi-agent coordination. Most benchmarks either focus on single-agent task completion or summarize multi-agent behavior with overall task success rates, which can obscure coordination failures such as duplicated work, violations of ordering constraints, resource contention, and desynchronized handoffs. In this paper, we introduce CoCoBench, a construct-level benchmark for evaluating multi-agent embodied coordination in executable household tasks. CoCoBench contains 897 oracle-validated instances organized around four recurring coordination constructs: task allocation, sequential ordering, mutual exclusion, and handoff coordination. In addition to task success rate, CoCoBench provides construct-level scores that measure whether agents coordinate effectively. We evaluate 11 leading MLLMs across different coordination modes, observation inputs, and numbers of agents. The results show that coordination ability is highly construct-specific: strong overall performance does not imply balanced competence across different coordination types. These findings point to new directions for designing targeted model architectures and improving multi-agent coordination ability.

View source

Similar papers

Preprint Sep 2026

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with...

Raphael Shu, Yu-Sen Zhang, Y. Cho et al. · 0 citations
Preprint Aug 2026

AgentRoom: Concurrent Multi-Agent Coding in a CRDT-Backed Shared Workspace

AgentRoom is a realtime collaborative editing protocol for concurrent coding agents that exposes file-level claim, status, and broadcast as MCP tools on a CRDT-merged shared filesystem and has less run-to-run variation than CLI-stable models.

Seonglae Cho, Donghyun Lee · 0 citations
Preprint Aug 2026

Bayesian Partner Modelling enables Adaptive Replanning for LLM Coordination

BayesBeliefAgent is introduced, which pairs a hierarchical LLM planner with a Bayesian tracking module and evaluates performance using replanning efficiency and the belief-action gap: the fraction of total decisions where an agent with a correct partner estimate executes a non-complementary skill.

Harsh Goel, A. S. Ellendula, Vaishnav Tadiparthi et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Rethinking Multi-Agent Collaboration: When More Is Less

The rapid advancement of large language models and single-agent harnesses has reshaped the landscape of autonomous systems, raising a critical question of when multi-agent collaboration offers genuine value. As individual agent capabilities continue to scale, multi-agent collaboration faces diminishing returns while in...

Yizhen Yuan, Yi-Bo Wu, Yi-Han Zhang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MASkills: Continual Skills Optimization for Multi-Agent LLM Systems

MASkills presents a new agent-optimization pipeline that integrates skill-conditioned credit assignment, hierarchical credit aggregation, and momentum-smoothed optimization, enabling agent skill libraries to evolve through refinement, induction, consolidation, and pruning.

Huaiyuan Yao, Xiaoou Liu, Charles Fleming et al. · 1 citation
Preprint Aug 2026

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

To test whether the taxonomy supports mitigation, TART, Taxonomy-Guided Actionable Representation, is introduced that makes the taxonomy's key aspects explicit to the planner and downstream sub-agents and consistently improves performance.

Vikas Pahuja, J. Brokman, O. Hofman et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.