Preprint
Aug 2026
Grounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test Generation
It is shown that test suites generated by the spec-driven agent are superior to the baseline and human-authored tests in 77.8% and 56.7% of the cases, respectively, and demonstrated improvements on following best practices, readability, and edge-case coverage.
Michele Tufano, James McClure, José Cambronero et al.
· 0 citations