Skip to content

Author

Pengfei Gao

We have 6 of 30 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Update from Hell: Can Coding Agents Survive Hidden Breakage in Dependency Upgrades?

Evaluating mainstream coding agents on DEPBENCH, a benchmark consisting of 203 real-world dependency-upgrade tasks across five package ecosystems spanning five language communities, each involving hidden code-level changes that require source code adaptation.

Zi-Jian Luo, Runzhi He, Peng-Fei Gao et al. · 0 citations
Preprint Jul 2026

Can Large Language Models Generate Observability-Aware Code?

Recent advances in coding agents have enabled the generation of increasingly complex software systems. While existing evaluations primarily focus on functional correctness, production systems must expose failure evidence to support observability. In this paper, we present a systematic study of observability in agent-ge...

Yongliang Tao, Hongyu Zhang, Pengfei Gao et al. · 0 citations
Preprint Aug 2026

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

The proposed AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches, suggests that automatic harness optimization is a promising path toward more performant and reliable agent...

Sungho Park, Wonjoong Kim, Rongyuan Tan et al. · 9 citations
Preprint Jul 2026

LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation

Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development. Existing benchmarks often center on localized tasks or end-state outcomes, offering limited insight into sustained execution. We introduce LOOPSBENCH, a...

Han Li, Zhemin Fang, Rili Feng et al. · 1 citation
#software testing Preprint Aug 2026

DPIAgent: Divide, Protocol, Isolate for Agentic Reproduction Test Generation

DPIAgent is proposed, a structured agentic framework built on three principles, Divide, Protocol, Isolate (DPI), that mitigates compound-objective ambiguity and goal drift, and shows that architectural structure and backbone capability are complementary axes rather than substitutes, demonstrating DPI's generalizability...

Hao Liu, Steven Liu, Xin Zhang et al. · 0 citations
Jul 2026

Change2Task: From Repository Changes to Executable Coding Agent Tasks and Environments

This work presents Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository, and provides executable data for coding agent training and evaluation while reducing repeated environment setup, storage, and task constructi...

Haomin Qi, Xing-Liang Wang, Xuanqi Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.