Vulnerability-detection benchmarks score the verdict an agent reaches, not the evidence it gathered. A model that recalls a CVE from pretraining therefore scores the same as one that traced the data flow. We study a task where this difference matters, deciding whether a commit introduces a vulnerability. Instead of sco...
Yi-Kun Li, Jin-Feng Jiang, Yuheng Yieh et al.· 0 citations
This work introduces Path2Spec, a divide-and-conquer framework that leverages LLMs to extract all execution paths from an input program, generates path-specific specifications for each, and merges them into a comprehensive overall specification.
Dan Huang, Zhensu Sun, Hui-Hui Huang et al.· 0 citations
Agent systems rely on LLM APIs for every response, but these APIs can return server errors, truncated responses, or corrupted content that propagates through downstream agents and causes task failure. Evaluating robustness under these faults is crucial for reliable deployment. Existing fault injection methods are offli...
Gou Tan, Zhensu Sun, Jieke Shi et al.· 0 citations
Real-world vulnerabilities often span multiple functions, yet most learning-based detectors classify each function in isolation: on a sample of real CVEs, we find that 71.7% of vulnerable functions require evidence from outside the function to be classified correctly. Agentic reinforcement learning (RL) could close thi...
Yikun Li, Ting Zhang, Jiakun Liu et al.· arXiv.org· 2 citations
Accurate vulnerability severity assessment is essential for prioritizing remediation, yet manually assessing Common Vulnerability Scoring System (CVSS) base metrics remains labor-intensive. Existing automated approaches often fail to capture the repository-level evidence required for assessing many CVSS base metrics. S...
Jinfeng Jiang, Yikun Li, Chengran Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.