Coding agents are increasingly asked to build software rather than patch it, yet benchmarks for from-scratch repository construction are mostly limited to a single language and depend on manually curated tasks. We introduce Zero2Repo, a benchmark in which an agent receives a product requirements document, an interface...
Pei Yang, Tian-Yu Shi, Yu-Hang Yao et al.· 0 citations
ShareMem is introduced, a memory architecture that shares reusable experience while grounding its application in the receiving user's own preferences, and improves step success, average task success, and dialogue-macro coding scores, respectively, over matched user-local memory across all four models.
Jinming Hu, Haodong Zhao, Qi Jia et al.· 0 citations
Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear trajectory from simple to complex semantic processing. While early AI systems mainly addressed tasks involving direct and literal semantic percep...
Xiujie Song, Ge-Fei Yang, Yi-Ning You et al.· arXiv.org· 0 citations
Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields. In this work, we introduce SciMIF, a novel benchmark designed to evaluate the capability of MLLMs in following complex...
Ye Shen, Yu-Ting Zheng, Dun Pei et al.· 0 citations
This paper introduces SafeAtlas-VL, a dataset of 1.5M training instances that places image-, request-, and response-level judgments on a five-level ordered scale, and trains the SafeAtlas Guard series of models via target-conditioned tuning for multimodal safety detection.
Zong-Rui Wang, Xiang-Yang Zhu, Sixiang Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.