FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs
Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the model to generate a proof-of-concept input that triggers a predefined target vulnerability. However, this setup may overlook valid crashes disc...