A Comprehensive Empirical Analysis of Patch Presence Testing: Capabilities, Limitations, and Paths Forward
Abstract
Patch presence testing determines whether a binary incorporates the security fixes associated with a known vulnerability and has become increasingly important in software supply-chain security. However, despite numerous proposed techniques, the field still lacks a comprehensive understanding of the practical capabilities and limitations of existing approaches. Fundamental questions remain unanswered, including how well current tools perform in real-world settings, which vulnerability or patch characteristics shape detection accuracy, and what underlying factors limit the effectiveness of all existing tools. To address these issues, we conduct the first systematic and in-depth empirical study of patch presence testing for C/C++ binaries. We construct a high-fidelity benchmark comprising 561 CVEs across ten widely used projects, with binaries compiled under diverse configurations. Using this dataset, we perform an extensive evaluation of five state-of-the-art tools representing both syntactic and semantic methodologies. Our findings show that: (1) accuracy reported in prior work reflects only cases where tools successfully generate outputs, whereas in practice many tools frequently fail to produce any result; (2) patch semantics, code scale, and compiler options exert a strong influence on accuracy, whereas CWE categories provide little predictive value; (3) common failures fall into two major categories: algorithmic limitations, such as the inability to detect subtle or evolved patches, and engineering deficiencies, such as failures triggered by function-level structural modifications or symbol duplication. Building on these findings, we develop two improvement strategies and integrate them into state-of-the-art tools, resulting in notable gains in both accuracy and overall reliability for patch detection.