LLMSuite is proposed, a hybrid test generation framework that integrates self-refinement prompting with class-level LLM reasoning into the search-based testing process and complements manually written test suites by exercising domain-specific behaviors that are often left untested.
Abstract
Automated unit test generation tools like EvoSuite perform well on general-purpose software but often struggle with domain-specific software such as Natural Language Processing (NLP) libraries, where inputs must follow semantic, syntactic, and structural constraints. Large Language Models (LLMs) can generate domain-relevant test code, but tests produced by LLMs alone often fail to compile or achieve sufficient coverage. We propose LLMSuite, a hybrid test generation framework that integrates self-refinement prompting with class-level LLM reasoning into the search-based testing process. In this mechanism, the LLM iteratively improves its test snippets based on feedback from previous generations. This enables the model to produce increasingly precise, domain-consistent code fragments that steer the evolutionary search toward exercising complex and otherwise hard-to-reach behaviors. When no objective improves over multiple generations in the underlying evolutionary algorithm, these refined snippets are parsed and injected into EvoSuite's population to expand the search space. To support our evaluation, we constructed a new dataset comprising 100 classes drawn from five widely used Java NLP projects. We also re-implemented CodaMOSA, a recent hybrid SBST-LLM technique, in Java to enable a direct comparison. Across this dataset, LLMSuite improves branch and line coverage by approximately 10% and 8%, respectively, and achieves an 11% higher mutation score than CodaMOSA-J. Compared to EvoSuite, LLMSuite yields roughly 15% higher branch and line coverage and 5% higher mutation score. Against an LLM-only baseline, it improves structural coverage by 36% and mutation score by about 24.7 percentage points. Finally, LLMSuite complements manually written test suites by exercising domain-specific behaviors that are often left untested.
GenCTL is proposed, a prompt-based framework for effective and model-aware NL2CTL translation without task-specific fine-tuning that improves the reliability and practical checkability of LLM-generated CTL specifications.
Ran Tao· Poster Volume 0007 The 2026...· 0 citations
Study of an end-to-end NL-to-PDDL pipeline that combines LLM generation, checks in terms of PDDL parsing, planning and validation, a domain-conformance checker, an LLM critic, and iterative repair shows that structured repair can be useful, and that PDDL~2.1 remains challenging for reference reconstruction, even when o...
J. Rosa, Pedro Santos, Valdemar Oliveira et al.· 0 citations
This paper presents an empirical evaluation of Large Language Models (LLMs) for automated model-based test generation, compared with a state-of-the-art model-based testing tool (GraphWalker) and its built-in algorithms (random and quick random for edge and vertex coverage settings).
H. Şanli, Onur Kilinççeker, Cihat Çetinkaya· 0 citations
XREPOTEST is introduced, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby, and Invocation Rate is proposed to assess whether generated tests meaningfully exercise the intended functionality.
L. Dung, Dong Cao Van, Nam Le Hai et al.· 1 citation
Two novel contributions are introduced: CodeEval and CodeQual, an open-source execution framework that provides researchers with a ready-to-use evaluation pipeline for evaluating and improving LLMs in software engineering contexts, encompassing both functional correctness assessment and subjective code quality evaluati...
Conventional automated REST API testing approaches often depend on rule-based logic, extensive configuration, or source-code access, which limits their adaptability in rapidly evolving development environments. Recent advances in Large Language Models (LLMs) offer new possibilities for automating API test generation th...
Phyu Thet Thet Kyaw· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.