GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents...
Dai-Feng Li, Huiqiang Jiang, Chengruidong Zhang et al.· 0 citations
Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requires diverse agentic environments, dependable outcome signals, and low extension cost. Existing generation pipelines commonly construct an environment before definin...
Xin-Jie Shen, Wei Fan, Xu-Dong Guo et al.· 0 citations
BabelFlow is introduced, a benchmark-general agentic workflow that adapts existing agent benchmarks to new languages by analyzing runtime dependencies, coordinating structure-preserving translation, and combining multi-layer verification with human review to preserve task and evaluation semantics.
Peng Kuang, Yu-Chun Fan, Jiang-Nan Li et al.· 0 citations
E-Commerce Bench is introduced, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events into a year-long business operation, and it is found that no single model dominates.
Wei Fan, Xin-Jie Shen, Xu-Dong Guo et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.