Conference
Open access
2026
AutoSUIT Bench - Automated Security UnIt Test Benchmark for LLM Coding
Upon benchmarking against LLMs, it is found that functionality pass rate is consistently higher than vulnerability pass rate for all programming languages, highlighting the necessity of vulnerable code benchmarks with larger CWE coverage.
Samuel Osebe, Fan Yang, Junyi Li et al.
· Annual Meeting of the Associ... · 0 citations