Repository-level issue-resolution benchmarks have made executable evaluation central to software-engineering agents, but their language coverage remains concentrated in mainstream ecosystems. Pony presents a different regime: it combines actors, reference capabilities, ahead-of-time compilation, and a rapidly evolving...
Bang Xie, Hao Liu, Zhen-Yu Shi et al.· 0 citations
AppEval is presented, a benchmark and native-toolchain evaluation framework for mobile application repair across HarmonyOS/ArkTS, iOS/Swift, and Android/Kotlin, and shows that mobile repair performance depends strongly on the evaluated agent while demonstrating why runtime-aware acceptance is necessary for meaningful c...
Bang Xie, Hao Liu, Zhen-Yu Shi et al.· 0 citations
OdinEval is presented, a reproducible benchmark built from documented defects in public Odin repositories built from documented defects in public Odin repositories, that evaluates six language models on 168 filtered instances under one shared protocol.
Bang Xie, Hao Liu, Zhi-Yuan Peng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.