Skip to content
Open access

Optimizing Context and Cost in LLM‐Based Unit Test Generation: A Study on External Dependency Retrieval Strategies

Aug 2026 · Expert systems · Vol 43 · 0 citations · 13 references

TL;DR

A systematic empirical study of multiple strategies for context enrichment and optimization in LLM‐based unit test generation, conducted on seven diverse projects (three open‐source and four proprietary industrial systems), encompassing 261 distinct methods establish this optimized context strategy as a cost‐effective solution for scalable, industrial‐grade automated test generation.

Abstract

While Large Language Models (LLMs) have demonstrated remarkable capabilities in code generation, their effectiveness in unit testing is often constrained by insufficient context regarding external dependencies. This limitation is particularly pronounced in industrial settings, where proprietary code remains opaque to the model. To address this challenge, we present a systematic empirical study of multiple strategies for context enrichment and optimization in LLM‐based unit test generation, conducted on seven diverse projects (three open‐source and four proprietary industrial systems), encompassing 261 distinct methods. By evaluating seven implementations (ranging from basic prompts to optimized context reduction strategies) across 10 independent runs, we analysed a total of 28,710 test suites. Our results demonstrate that combining prompt engineering with external dependency retrieval achieves an average branch coverage increase of 11.52 percentage points on industrial software over the baseline, with statistically significant improvements across all competing implementations. Beyond coverage, richer context substantially reduces generation‐repair iterations, cutting median execution time by 51.3% in industrial projects. We further show that reducing external dependencies to method signatures alone decreases input token consumption by up to 46.6% (25.4% in industrial projects) while fully preserving the coverage and efficiency gains of the complete retrieval approach. To confirm that these benefits are not tied to a specific model, we replicate the core comparison across three LLM backends from different families, obtaining a consistent, statistically significant coverage improvement on industrial code in every case. These findings establish this optimized context strategy as a cost‐effective solution for scalable, industrial‐grade automated test generation.

Read PDF

Similar papers

#software testing Preprint Aug 2026

XREPOTEST: Benchmarking Multilingual Repository-Level Unit Test Generation for Large Language Models

XREPOTEST is introduced, a multilingual repository-level benchmark for unit test generation spanning five underexplored languages: Rust, Go, Julia, PHP, and Ruby, and Invocation Rate is proposed to assess whether generated tests meaningfully exercise the intended functionality.

L. Dung, Dong Cao Van, Nam Le Hai et al. · 1 citation
Conference Jul 2026

Improving LLM-Based Unit Test Generation Through Root-Cause-Driven Prompt Design

This paper addresses automated unit test generation with large language models (LLMs). LLM-based test generation has not yet attained a quality level sufficient for practical use in industry. Although LLMs often reproduce API syntax faithfully, they frequently disregard semantic usage constraints and execution-environm...

Mizuki Yamada, Masahiko Kato, Juichi Takahashi · 0 citations
Preprint Aug 2026

Doc2CI: A Multi-Service Study of CI Configuration Generation Using Large Language Models

A large empirical study on using LLMs to generate CI configurations from natural language across services and model families suggests that similarity and validity are distinct objectives for CI generation and motivate schema-aware evaluation and tooling for LLM-based configuration generation.

T. A. Ghaleb · 0 citations
Open access

Evaluation and Distillation of Source Code Generation Tasks by Large Language Models

Two novel contributions are introduced: CodeEval and CodeQual, an open-source execution framework that provides researchers with a ready-to-use evaluation pipeline for evaluating and improving LLMs in software engineering contexts, encompassing both functional correctness assessment and subjective code quality evaluati...

Danny Brahman · 0 citations
#software testing Preprint Sep 2026

Compound Prompt Constraints in LLM Code Generation: A Factorial Study of Format, Persona, and Urgency

Large language models (LLMs) are increasingly used in software engineering pipelines for code generation, where production prompts often combine multiple constraints. This paper presents a full-factorial empirical study of how output formatting, persona assignment, and urgency framing jointly affect LLM code-generation...

Shrenik Jadhav, Nickalsa LaPlaca, Caleb Stone et al. · 0 citations

Capacity vs. architecture: an evaluation of SLMs for automated docstring generation

A reproducible, human-validated evaluation framework applied to 13 strategies—four architectural families crossed with four reasoning variants crossed with four reasoning variants—across three SLMs spanning 3B–14B parameters, plus targeted ablations.

Balaji Venktesh, Amsaprabhaa M, G. Sundaram · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.