Skip to content

VEX-Bench: Benchmarking LLM Agents for Assessing Exploitability of Software Supply Chain Vulnerabilities

Sep 2026 · 0 citations · 41 references
Computer Science

TL;DR

VEX-Bench is introduced, the first benchmark for evaluating LLM agents'ability to assess the exploitability of software supply chain vulnerabilities, and contains 75 real-world cases mined from GitHub and labeled by security experts, covering Python, Java, and Go.

Abstract

The software supply chain has become an increasingly exposed attack surface because of its reliance on intricate yet fragile dependencies. Existing defenses such as GitHub Dependabot often raise many false alerts because their coarse-grained matching cannot determine whether a vulnerable dependency is actually exploitable. Security analysts typically spend substantial time assessing vulnerability exploitability case by case. Recent LLM agents have emerged as promising candidates for this task given their advanced capabilities in coding and cybersecurity, yet no existing benchmark evaluates them on it. Prior benchmarks target zero-day settings, where agents detect and exploit previously unknown vulnerabilities. In contrast, software supply chain security focuses on how known vulnerabilities in upstream dependencies affect downstream projects. This requires agents to reason across repositories and determine whether an upstream vulnerability is exploitable in the downstream project. To address this gap, we introduce VEX-Bench, the first benchmark for evaluating LLM agents'ability to assess the exploitability of software supply chain vulnerabilities. It contains 75 real-world cases mined from GitHub and labeled by security experts, covering Python, Java, and Go. We evaluate nine models across three agent harnesses. While GPT-5.5 and Claude Opus 4.6 reach approximately 80% F1 on binary vulnerability-status classification, only GPT-5.5 surpasses 70% macro-F1 on fine-grained justification classification. This gap highlights the challenge of moving beyond binary exploitability assessment to identifying fine-grained exploitability reasons. Code and data: https://github.com/steven1518/vex-bench

View source

Similar papers

Preprint Aug 2026

CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training

CyberForge is a framework that synthesizes executable, repository-level security training data by injecting vulnerabilities into real C/C++ projects and holds 1034 validated vulnerabilities across 80 projects and 63 weakness categories, with edit locality similar to real CVE patches under a real-versus-real noise floor...

Amine Lbath, Manan Suri, A. Delaitre et al. · 0 citations
Preprint Sep 2026

ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch

Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an execu...

Liang He, Sheng Wu, Hao-Miao Hao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale

VLoc Benchmark results establish vulnerability localization as a distinct repository-scale capability and provide a setting for studying both how security agents search for vulnerable code and when they should refrain from reporting it.

Aman Priyanshu, Supriti Vijay, Kimia Majd et al. · 0 citations
#artificial intelligence Preprint Sep 2026

CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance

Large language model agents have demonstrated promising capabilities in cybersecurity tasks, yet their ability to reconstruct complete Advanced Persistent Threat attack campaigns from complex security logs remains largely unexplored. Existing cybersecurity benchmarks for agents mainly focus on vulnerability discovery,...

Qi Chen, Fu-Shuo Huo, Hang-Li Shen et al. · 0 citations
Open access Oct 2026

Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM Agents

LLM agents that invoke external tools face critical safety vulnerabilities when malicious manipulations exploit their implicit trust in tool outputs and metadata. However, identifying these vulnerabilities through testing is challenging due to the need to bypass safety guardrails with semantically legitimate inputs, th...

Yu-Chen Shao, Zi-Qun Bao, Yu-Heng Huang et al. · 0 citations
Preprint Sep 2026

Measuring the Security of the Evolving Software Supply Chain: a Research Agenda

Software supply chain security has become increasingly critical due to the widespread reliance on third-party dependencies and the growing attack surface of modern software ecosystems. However, existing quantitative, measurement-based analysis and vulnerability management approaches remain largely fragmented and ecosys...

Sarah Meriem Ourari · 0 citations

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.