Skip to content

MCPGen: Benchmarking LLMs on Executable MCPWorkflow Development

Sep 2026 · 0 citations · 37 references
Computer Science

TL;DR

MCPGen is introduced, an executable benchmark for Model Context Protocol (MCP) workflow development that evaluates three diagnostic tasks: workflow reconstruction, tool creation, and backward-compatible workflow extension and evaluates 11 representative LLMs in a single-turn foundation-model setting.

Abstract

We study whether LLMs can produce executable workflow artifacts that remain consistent across graph structure, tool implementation, schema bindings, and runtime wiring. In this setting, correctness depends on cross-layer consistency: a workflow may be structurally plausible, yet still fail because tool implementations, schema bindings, or runtime execution do not align. Existing benchmarks largely evaluate these capabilities in isolation or rely on trajectory-level proxies, leaving open whether generated workflow artifacts execute end-to-end. We introduce \textbf{MCPGen}, an executable benchmark for Model Context Protocol (MCP) workflow development. MCPGen contains 100 self-contained MCP projects across 16 application domains and evaluates three diagnostic tasks: workflow reconstruction, tool creation, and backward-compatible workflow extension. We evaluate 11 representative LLMs in a single-turn foundation-model setting, assessing generated artifacts through static analysis, unit and integration tests, and process-isolated end-to-end execution. Models reach 88.5\% on workflow reconstruction, but no model exceeds 57\% end-to-end execution success. Per-tool unit-test pass rates reach 63.8\%, while project-level integration success does not exceed 45\%, suggesting that integration remains a major bottleneck even when isolated tool tests pass.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch

Coding agents powered by large language models (LLMs) are evolving from making localized code changes to developing complete software repositories. However, evaluating repository-scale generation remains challenging: tasks must demand system-level reasoning while ensuring that all evaluated behaviors are precisely spec...

Hantian Ding, Chloe Bi, Jia-Cheng Zhu et al. · 0 citations
Preprint Sep 2026

Semantics, Workflows, and Infrastructure: Understanding Agent Serving at Production Scale

Large language model (LLM) agents execute applications through a workflow of inference requests with tool calls and user interactions. Serving these applications at production scale requires understanding how application behavior shapes inference demand and for guiding efficient execution. Recent characterization studi...

Yi-Hao Zheng, Jing-Zhe Jiang, De-Jiang Zhu et al. · 0 citations
#machine learning Preprint Sep 2026

DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis

A structurally valid DeFi workflow can still authorize a costly trade. We introduce DeFiFlowBench, a benchmark of 207 team-authored prompts for natural-language DeFi workflow synthesis. It measures graph coverage, configuration completeness, and declared safety predicates, then tests supported trade configurations on a...

Abhinav Kumar, Harshit Arora, Varun Singh et al. · 0 citations
Book Open access Oct 2026

Beyond Single-run Correctness: Nondeterminism-aware Evaluation of LLM-based Model Transformations

Model transformation is a core model-driven engineering (MDE) operation in which reproducibility is expected: under fixed metamodels, source model, and transformation rules, a deterministic engine should produce a stable target model. Large Language Models (LLMs) are increasingly explored for MDE tasks, but evaluations...

Riccardo Rubei, Alessio Bucaioni, A. Di Salle · 0 citations
Preprint Aug 2026

DepWareTrans: Dependency-Aware Incremental Repository Migration across Co-executable Languages

This paper proposes a dependency-aware incremental migration framework that elevates the unit of translation from individual files to dependency-consistent batches and improves scalability and reliability in repository-level code translation.

Sivajeet Chand, Alexander Pretschner, Steve Haupt et al. · 2 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.