Skip to content

SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction

Sep 2026 · 0 citations · 13 references
Computer Science

TL;DR

This work recasts vulnerability discovery as an input-prediction task with a closed, deterministic ground truth, and decomposes discovery into three task modes over 22 real-world C/C++ programs spanning 15 domains, finding constraint inference, not navigation, is the dominant bottleneck.

Abstract

Vulnerability discovery is becoming an important ability of large language model (LLM) agents: agents that silently miss real defects leave critical software exposed. Rigorously measuring this ability is therefore urgent, but existing benchmarks are gameable through data contamination, score recall against an unknowable vulnerability set, often rely on synthetic bugs, and report a single end-to-end verdict that cannot localize where an agent fails. Vulnerability discovery is a composite ability: an agent must comprehend source code, infer input constraints, construct inputs, execute them, and iteratively correct from feedback. We recast its measurement as an input-prediction task with a closed, deterministic ground truth: using coverage-guided fuzzing, we mine deep target branches in real-world C/C++ programs and ask an agent to predict an input that drives execution to a given branch. This decomposes discovery into three task modes over 22 real-world C/C++ programs spanning 15 domains. Open-loop and Feedback-enabled share 60 fixed-target task instances across 16 of these codebases (13 domains), testing input construction without and with a distance oracle to isolate code comprehension from feedback-driven correction. Online Arena instead removes the predefined target and scores path exploration by coverage gain on a separate, partially overlapping pool of 11 programs; agents collectively confirmed 13 distinct bugs across six programs. Evaluating 15 default-effort model-scaffold configurations, the best reaches only 55.0% pass rate in the Feedback-enabled mode, and the mean across seven paired Claude Code configurations is 36.4% with feedback versus 19.3% without. Decomposing failures, we find constraint inference, not navigation, is the dominant bottleneck. We release SWE-Test with a turnkey evaluation environment.

View source

Similar papers

Preprint Aug 2026

FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs

Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the model to generate a proof-of-concept input that triggers a predefined target vulnerability. However, this setup may overlook valid crashes disc...

Ze Sheng, Aleksandar Kezic, Zhicheng Chen et al. · 0 citations
Preprint Sep 2026

Reveree: Diagnosing LLM Reverse-Engineering Agents

Reveree is a diagnostic framework that scores an LLM RE agent's trajectory at three tiers: solve rate, milestone progress through an eight-stage RE schema, and a behavioral profile of its actions, finding that the base model dominates performance, whereas prompting strategy is a secondary, model-dependent effect.

Hadjer Benkraouda, Hongyu Cai, Berkay Celik et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations

LLM-based code generation is now embedded in mission-critical pipelines, but defenses against vulnerable output remain post-hoc -- static analyzers, fine-tuned classifiers, or an LLM judge that screen completed code, ignoring the generating model's own internal state. We test a narrower, directly measurable question: w...

Alizishaan Khatri · 0 citations
Preprint Sep 2026

ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch

Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an execu...

Liang He, Sheng Wu, Hao-Miao Hao et al. · 0 citations
Open access 2026

From Trace to Line: An Empirical Study of What Drives LLM-Based OSS Vulnerability Localization

This paper introduces T2L (Trace-to-Line), a reproducible research framework that narrows repository-scale code into candidate vulnerable lines through AST-based chunking, structured diagnostic information collection, and evidence-guided refinement that improves trace-to-line localization.

Hao-Ran Xi, Ming-Hao Shao, Brendan Dolan-Gavitt et al. · 0 citations
#machine learning Preprint Sep 2026

Introspective Uncertainty Estimation for LLM-Based Code Generation

The findings suggest that hidden states are a robust and informative resource for estimating functional code correctness, supporting a two-stage workflow that combines response-level risk screening with targeted line-level prioritization.

T. Klassert · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.