STARS: Static Analysis-Guided Assertion Synthesis using Large Language Models
Abstract
Automated unit test generation promises to reduce the cost of software quality assurance, and hence, is attracting attention from both academia and industry. Yet, generating assertions that are executable, meaningful to developers, and able to catch faults remains an unsolved challenge. Existing approaches either randomly enumerate assertions that are plausible based on static program analysis without considering whether they naturally fit the test prefix or query an LLM to generate assertions based on local context only, such as the test prefix and the focal method. However, we observe that local context alone is insufficient for LLMs to generate high-quality assertions because many desirable assertions are built from components that are almost impossible to guess for an LLM, such as sequences of multiple method calls. This paper presents STARS, a novel test assertion generation technique that combines the benefits of static program analysis and LLM-based synthesis. The key idea is to first gather a set of assertion components based on static program analysis and to then combine, concretize, prioritize, and improve them with an LLM. The resulting assertions go beyond what an LLM alone could realistically guess based on the test prefix and focal method, and they naturally fit the given test case. Empirical results show that STARS consistently outperforms the state-of-the-art baseline in five evaluation metrics. STARS achieves an exact-match rate of at most 83.4% using GPT-5.4. Compared with the baseline, STARS’s mutation score nearly triples that of the baseline (10.97% vs. 3.97%), approaching that of developer-written assertions, while consuming 24.75% fewer tokens and 30.6% fewer LLM queries.