Preprint
Jul 2026
SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
SQBench is introduced, a benchmark for evaluating production-oriented task delivery by language-model agents and shows that functional completion alone does not fully characterize delivery quality and that risk determinations should be reported separately.
Sum Sun
· 0 citations