MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
A novel benchmark for evaluating the task-solving capabilities of LLM agents under dynamic toolset evolution, and proposes 11 mutation operators to simulate realistic tool evolution within 123 MCP servers, establishing MCPEvol-Bench as a standard for evaluating agent adaptability in dynamic tool environments.