This work proposes ExplainBench, a benchmark to automatically evaluate explanations from coding agents, based on the intuition that informative explanations should enable an LLM to correctly answer questions, allowing quantitative comparison of explanation quality between agents.
Zhiyuan Pan, Sungmin Kang, Imam Nur Bani Yusuf et al.· 0 citations
This paper presents the experience and lessons learned in adapting the AutoCodeRover program improvement agent to automatically propose patches for issues reported by SonarQube, and names this new agent SonarQube Remediation Agent, specialized for fixing SonarQube issues.
Martin Mirchev, Ridwan Shariffdeen, Haifeng Ruan et al.· SIGSOFT FSE Companion· 1 citation
It is argued that risk-free deployment must be grounded in the agent's trajectory: the recorded sequence of reasoning steps, tool invocations, and environmental observations, and the absence of adequacy metrics.