Preprint
Jul 2026
ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance
ChainSWE is introduced, the first benchmark for evaluating agents on sequential, dependent bug fixes within a shared codebase, and reveals a consistent performance drop by up to 70% as the chain length increases.
Qirui Jin, Lingching Tung, Kenan Li et al.
· 1 citation
· ⚡1