The goal is to understand the code generation errors of foundation LLMs and explore the solution to resolve directly fixable errors, and to design and evaluate the LlmFix fixing method and constructed the LlmErrorEval dataset.
Hao Wen, Yue-Heng Zhu, Chao Liu et al.· Empirical Software Engineeri...· 0 citations
Findings show that functional-only evaluation overestimates agents'ability to satisfy the full requirements of repository-level repair tasks, and introduces SWE-Gate, a repository-level benchmark for software engineering agents that explicitly evaluates review constraint compliance alongside functional correctness.
Xin He, Yan-Lin Wang, Ming-Wei Liu et al.· 0 citations
MCR-Bench is introduced, the first defect state-aware benchmark designed for realistic multi-round code review, and in-depth error analysis dissects the distinct drivers of false positives and false negatives, revealing critical weaknesses such as cross-round temporal misalignment and inadequate long-range memory.
De-Wu Zheng, Yan-Lin Wang, Xi-Wen Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.