Evaluating mainstream coding agents on DEPBENCH, a benchmark consisting of 203 real-world dependency-upgrade tasks across five package ecosystems spanning five language communities, each involving hidden code-level changes that require source code adaptation.
Zi-Jian Luo, Runzhi He, Peng-Fei Gao et al.· 0 citations
Recent advances in coding agents have enabled the generation of increasingly complex software systems. While existing evaluations primarily focus on functional correctness, production systems must expose failure evidence to support observability. In this paper, we present a systematic study of observability in agent-ge...
Yongliang Tao, Hongyu Zhang, Pengfei Gao et al.· 0 citations
The proposed AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches, suggests that automatic harness optimization is a promising path toward more performant and reliable agent...
Sungho Park, Wonjoong Kim, Rongyuan Tan et al.· 9 citations
Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing benchmarks often center on localized tasks or end-state outcomes, offering limited insight into sustained execution. We introduce LOOPSBENCH, a...
DPIAgent is proposed, a structured agentic framework built on three principles, Divide, Protocol, Isolate (DPI), that mitigates compound-objective ambiguity and goal drift, and shows that architectural structure and backbone capability are complementary axes rather than substitutes, demonstrating DPI's generalizability...
Hao Liu, Steven Liu, Xin Zhang et al.· 0 citations
This work presents Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository, and provides executable data for coding agent training and evaluation while reducing repeated environment setup, storage, and task constructi...