Aug 2026· Message Understanding Conference· 0 citations· 9 references
Computer Science
TL;DR
AlHAZEN and AVICENNA show that effective debugging tools have tradeoffs between accuracy and interpretability to support developers’ decision-making in increasingly complex software environments.
Abstract
Debugging is an important activity in software development, yet providing actionable and comprehensible explanations for program failures remains challenging. Automated tools such as ALHAZEN and AVICENNA address this by using distinct strategies: ALHAZEN employs binary decision trees to show failure-inducing conditions, while AVICENNA uses a specification language to model complex input dependencies. To assess their impact on usability and user efficiency, we conducted a controlled within-between-subjects user study with 18 participants tasked with resolving four software bugs using either tool or no support. Quantitative results showed that both tools improved debugging efficiency compared to manual methods, with AVICENNA offering more precise diagnostics but requiring higher cognitive effort. Qualitative feedback revealed a preference for AVICENNA’s expressiveness despite its complexity. Our findings show that effective debugging tools have tradeoffs between accuracy and interpretability to support developers’ decision-making in increasingly complex software environments.
Goanna is introduced, a novel type checker for Haskell that focuses on improving error diagnostics, and shows performance constraints when diagnosing large programs containing complex errors, but remains responsive enough to provide real-time debugging assistance for small to medium-sized programs.
Shuai Fu, Tim Dwyer, Peter James Stuckey et al.· International Conference on...· 0 citations
While LLM-rewritten messages significantly improved subjective evaluations, with pragmatic messages rated as clearer and less cognitively demanding, these perceived gains did not translate into statistically significant improvements in objective debugging performance.
Alexandru-Radu Moraru, Shreyan Biswas, U. Gadiraju· 0 citations
Large language models (LLMs) have shown promise for automated program repair, but it remains unclear which debugging signals are most useful and when additional context becomes distracting, costly, or ineffective. We present a controlled empirical study of LLM-based bug fixing on FIXEVAL, comparing three model families...
Dinesh Kumar Gummadavelli, Xian-Shan Qu, Xiao-Peng Li et al.· EAI Endorsed Transactions on...· 0 citations
An empirical study involving 5 Large Language Models and 4 benchmarks evaluates the effectiveness and efficiency of 3 widely used adequacy criteria: statement coverage, branch coverage, and mutation testing, finding that mutation testing only marginally outperforms traditional coverage criteria in both triggering and d...
Asma Hamidi, Michael Konstantinou, R. Degiovanni et al.· 0 citations
This work proposes tHinter, an automated approach that frames translation error localization as a differential testing task, and integrates mixed-factorial user studies, expert validation, and SWOT-based strategic analysis to assess the perceived helpfulness and resilience within the rapidly evolving LLM landscape.
Shengnan Wu, Xin-Yu Sun, Xin Wang et al.· ACM Transactions on Software...· 0 citations
This study empirically evaluates whether cost-efficient Large Language Models (LLMs) can be trusted to generate enterprise code to a written specification. Three models (Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5) were asked to solve 992 algorithmic problems as Java Spring Boot service methods conforming to a ma...
Chandimal Adikari, Nandika Herath· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.