Modern coding agents can deliver increasingly large repository-level changes, and recent benchmarks reflect this by emphasizing long-horizon tasks with large reference implementations. Many benchmarks evaluate coding agents'implementation capability to produce correct code edits from detailed specifications. However, p...
Yun Peng, Zi-Han Wu, Ze-Yang Zhuang et al.· 0 citations
This research provides novel insights for the development of highly active and low-toxic phenamacril derivatives, and lays the foundation for achieving a closed-loop research and development process involving structure-activity relationships, toxicity prediction, iterative optimization, and re-evaluation.
Yan-Ru Chen, Xiao-Rong Zhang, Yu-Hao Shi et al.· Pest Management Science· 0 citations
TraceProbe is presented, a trajectory-diagnostic framework that recovers what resolve rate hides and adds auditable diagnostic context to outcomes by localizing inspection targets, suggesting failure hypotheses, and prioritizing runs for review.
Rui Shu, C. Chong, Xin Zhou et al.· arXiv.org· 1 citation
SWE-RPG is introduced, a repository-level benchmark that combines executable patch evaluation with validated ground-truth references (GTs) for Requirement Clarification and Implementation Planning, and suggests implicit-requirement recovery as a key candidate direction for improving coding agents.
Xin Zhou, Chun-Yong Chong, Kisub Kim et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.