#artificial intelligence
Dec 2025
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
A benchmark based on real-world MCP definitions designed to evaluate the tool-use capabilities of agents, which reveals significant performance differences in handling complex, multi-step tool invocations.
Zixiang Liu, Wenrui Liu, Elsie Dai et al.
· arXiv.org · 10 citations