Skip to content

Rethinking the Evaluation of Efficiency Methods for Multi-Agent Systems

Sep 2026 · 1 citation · 36 references
Computer Science

TL;DR

This work introduces a controlled and MAS-demanding diagnostic benchmark for representative MAS efficiency methods and shows that many reported gains are setup-dependent and may arise from structural collapse, disabled tool pathways, or starting systems where random pruning already preserves accuracy, rather than robust improvements in MAS efficiency.

Abstract

Efficiency is increasingly important for Large Language Model (LLM)-based multi-agent systems (MAS), as larger models and more agents introduce substantial execution costs. Recent methods aim to make MAS cheaper by pruning agents, removing communication edges, or searching for compact structures. However, we argue that existing evaluations may overestimate their true ability to improve MAS efficiency. Reported gains are often measured under method-specific prompts and starting topologies, making them difficult to attribute to the proposed structural changes. Moreover, many reported successes appear in non-MAS-demanding settings, where a single agent or a randomly pruned system can already preserve strong performance. To study these issues, we introduce a controlled and MAS-demanding diagnostic benchmark for representative MAS efficiency methods. We evaluate methods under a shared backbone model, agent registry, and runtime, across controlled variations in topology, scale, depth, and tool use. Our analysis shows that many reported gains are setup-dependent and may arise from structural collapse, disabled tool pathways, or starting systems where random pruning already preserves accuracy, rather than robust improvements in MAS efficiency.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Rethinking Multi-Agent Collaboration: When More Is Less

SAIGE, a lightweight multi-agent collaboration mechanism based on Semantic-Aware Incremental Graph Evolution, is proposed, suggesting that multi-agent superiority is bounded by task structure rather than universal, and that more agents do not necessarily make a system more intelligent.

Yizhen Yuan, Yi-Bo Wu, Yi-Han Zhang et al. · 0 citations
Preprint Aug 2026

ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration

This work introduces a generalizable evaluation framework that maps native MAS traces into a shared space of unified collaboration graphs, enabling different methods to be evaluated under the same representation, reference set, and metric panel.

Guo Chen, Ziwen Li, Reed Li et al. · 0 citations
Preprint Sep 2026

MASTraceBench: Diagnosing Collaboration Gains through Proposal Trajectories in LLM-Based Multi-Agent Systems

LLM-based multi-agent systems (MAS) have shown promise in complex problem solving. As MAS methods diversify, systematic evaluation becomes increasingly challenging. However, existing benchmarks largely focus on final outcomes, leaving unclear how collaboration gains arise, are preserved, or are lost. To address this li...

Ya-Peng Li, Song-Ze Li, Shuang Yu et al. · 0 citations
Preprint Sep 2026

AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs

Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with...

Raphael Shu, Yu-Sen Zhang, Y. Cho et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Agentic Multi-Turn Reasoning: A Fairness Approach

Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where su...

Thanh-Dat Truong, Sankalp Pandey, Hugh Churchill et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.