SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction
This work recasts vulnerability discovery as an input-prediction task with a closed, deterministic ground truth, and decomposes discovery into three task modes over 22 real-world C/C++ programs spanning 15 domains, finding constraint inference, not navigation, is the dominant bottleneck.
Yuan-Xiang Shi, Jia-Yi Lin, Xuan-Yong Lin et al.
· 0 citations