MCPGen is introduced, an executable benchmark for Model Context Protocol (MCP) workflow development that evaluates three diagnostic tasks: workflow reconstruction, tool creation, and backward-compatible workflow extension and evaluates 11 representative LLMs in a single-turn foundation-model setting.
Yingxuan Yang, Jia-Qi Liu, Li-Rui Guan et al.· 0 citations
Long-term memory is essential for LLM-based agents operating over extended interactions. Existing memory systems primarily update memory when new information arrives, treating retrieval as the endpoint of memory access rather than a driver of memory evolution. Consequently, retrieval feedback is rarely exploited to reo...
Yuan-Yi Song, Yukai Wang, Xinbei Ma et al.· 0 citations
The key observation is that although expert trajectories are scarce, high-quality final artifacts such as literature reviews, analyst reports and legal judgments, are abundant in pre-training data and can be viewed as compressed traces of the evidence-seeking processes that produced them.
Junjie Huang, Jiarui Qin, Di Yin et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.