Open access
Jul 2026
PROBE: Benchmarking code generation in large language models
The findings show that, while LLMs achieve promising results, they struggle with harder problems and with programming languages that have fewer available resources for training, and they often fail due to fundamental and easily avoidable errors that underscore the unreliability of automatically generated code.
Rodrigo Pato Nogueira, Marco Vieira, João R. Campos
· Empirical Software Engineeri... · 1 citation