Coding-agent evaluation has progressed from resolving individual issues to carrying out long-horizon development, yet task completion is still largely assessed within a single codebase. In software ecosystems, many features and bug fixes require coordinated changes across multiple repositories. We introduce WideSWE to...
Bao-Yi Wang, Xing-Liang Wang, Jin-Yang Wu et al.· 0 citations
EviRCA is presented, a framework for LLM-based RCA that decouples deterministic evidence extraction from LLM reasoning and substantially outperforms prior OpenRCA baselines that achieve up to 15.2%, while reducing token consumption by 15-26x and execution time by 3-20x.
Yu-Hao Wang, Zhen Qin, Xing-Liang Wang et al.· 0 citations
This work presents Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository, and provides executable data for coding agent training and evaluation while reducing repeated environment setup, storage, and task constructi...