Semantic Tree: Leveraging Semantic UI Representation for Test Case Generation
Abstract
Software testing is a crucial process in software development. Despite its importance, the software testing process is often time-consuming and inconsistent when performed manually. With the rapid development of Large Language Models (LLMs), software testing processes such as generating test cases have utilized LLMs. Nevertheless, LLM performance often remains inconsistent for generating test cases. Therefore, this study evaluated the impact of prompting strategies and input representation on the performance of Large Language Models (LLMs) in generating test cases. Furthermore, this study aims to contribute more broadly to SDG 9 through innovation in software testing. This study proposed a representation transformation from the DOM to a Semantic Tree to reduce redundant information while preserving essential UI elements. We compared zero-shot and few-shot prompting strategies across three input representations: raw HTML, clean HTML, and Semantic Trees. The results indicated that the Semantic Tree representation generally achieved the highest average performance while consistently improving token efficiency compared to other representations. Statistical significance analysis further showed that the most influential factor were associated with the choice of LLM provider rather than input representation and prompting strategy. These findings suggest that both selecting an appropriate LLM and using efficient input representations are important for improving the practicality and scalability of LLM-based test case generation.