Skip to content
Book Open access

Does Great Power Come with Great Explainability? Comparing Explanation Strategies for Automated Program Diagnosis

Aug 2026 · Message Understanding Conference · 0 citations · 9 references
Computer Science

TL;DR

AlHAZEN and AVICENNA show that effective debugging tools have tradeoffs between accuracy and interpretability to support developers’ decision-making in increasingly complex software environments.

Abstract

Debugging is an important activity in software development, yet providing actionable and comprehensible explanations for program failures remains challenging. Automated tools such as ALHAZEN and AVICENNA address this by using distinct strategies: ALHAZEN employs binary decision trees to show failure-inducing conditions, while AVICENNA uses a specification language to model complex input dependencies. To assess their impact on usability and user efficiency, we conducted a controlled within-between-subjects user study with 18 participants tasked with resolving four software bugs using either tool or no support. Quantitative results showed that both tools improved debugging efficiency compared to manual methods, with AVICENNA offering more precise diagnostics but requiring higher cognitive effort. Qualitative feedback revealed a preference for AVICENNA’s expressiveness despite its complexity. Our findings show that effective debugging tools have tradeoffs between accuracy and interpretability to support developers’ decision-making in increasingly complex software environments.

Read PDF

Similar papers

Open access Sep 2026

Goanna: a novel approach for automated type error debugging

Goanna is introduced, a novel type checker for Haskell that focuses on improving error diagnostics, and shows performance constraints when diagnosing large programs containing complex errors, but remains responsive enough to provide real-time debugging assistance for small to medium-sized programs.

Shuai Fu, Tim Dwyer, Peter James Stuckey et al. · 0 citations
Preprint Aug 2026

Beyond the Traceback: Using LLMs for Adaptive Explanations of Programming Errors

While LLM-rewritten messages significantly improved subjective evaluations, with pragmatic messages rated as clearer and less cognitively demanding, these perceived gains did not translate into statistically significant improvements in objective debugging performance.

Alexandru-Radu Moraru, Shreyan Biswas, U. Gadiraju · 0 citations
Open access Sep 2026

When Does Context Help? A Controlled Study ofLLM-Based Bug Fixing

Large language models (LLMs) have shown promise for automated program repair, but it remains unclear which debugging signals are most useful and when additional context becomes distracting, costly, or ineffective. We present a controlled empirical study of LLM-based bug fixing on FIXEVAL, comparing three model families...

Dinesh Kumar Gummadavelli, Xian-Shan Qu, Xiao-Peng Li et al. · 0 citations
#software testing Preprint Sep 2026

How effective are traditional test criteria at detecting bugs in large language models generated code?

An empirical study involving 5 Large Language Models and 4 benchmarks evaluates the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing, finding that mutation testing only marginally outperforms traditional coverage criteria in both triggering and d...

Asma Hamidi, Michael Konstantinou, R. Degiovanni et al. · 0 citations
#software testing Review Sep 2026

Debugging Functionality-Twisting Translations by LLMs via Differential Testing with Bayesian Prior

This work proposes tHinter, an automated approach that frames translation error localization as a differential testing task, and integrates mixed-factorial user studies, expert validation, and SWOT-based strategic analysis to assess the perceived helpfulness and resilience within the rapidly evolving LLM landscape.

Shengnan Wu, Xin-Yu Sun, Xin Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

An Empirical Evaluation of Cost-Efficient Large Language Models on Algorithmic Programming Tasks

This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5) were asked to solve 992 algorithmic problems as Java Spring Boot service methods conforming to a ma...

Chandimal Adikari, Nandika Herath · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.