Preprint
Jul 2026
MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
A novel benchmark for evaluating the task-solving capabilities of LLM agents under dynamic toolset evolution, and proposes 11 mutation operators to simulate realistic tool evolution within 123 MCP servers, establishing MCPEvol-Bench as a standard for evaluating agent adaptability in dynamic tool environments.
Huanxi Liu, Kun Hu, Jiaqi Liao et al.
· 0 citations