PPTBench establishes a measurable testbed for studying visual coding and advancing agents toward more reliable visual creation, and shows that agents can generally produce valid slide files, but struggle to produce high-quality reconstructions that faithfully recover the semantics and visual structure of the target.
Xiao-Qiu Wang, Yi-Zhe Chi, Wen-Yi Li et al.· 0 citations
This work presents AI4AI\mbox{-}Bench, 10 frozen research repositories spanning 10 training algorithm families, and releases the task suite, the evaluators and every scored submission, so that the measurement can be repeated as these systems change.
Yi-Zhe Chi, Wenyi Li, De-Yao Hong et al.· 5 citations
SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.
De-Yao Hong, Yi-Zhe Chi, Wen-Yi Li et al.· 3 citations· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.