Skip to content

SyzHarness: Patch-Based Kernel Bug Reproduction with LLM-Synthesized Fuzzing Harnesses

Sep 2026 · 0 citations · 61 references
Computer Science

TL;DR

SyzHarness is a framework that combines LLM reasoning with coverage-guided fuzzing for patch-based Linux kernel vulnerability reproduction and achieves a 73% bug reproduction success rate, substantially outperforming prior directed greybox fuzzing.

Abstract

Automated kernel vulnerability reproduction is essential for bug triage, patch validation, and regression testing, but still lacks an effective and efficient solution. The core challenge is twofold: a reproducer must first recover the trigger scaffold needed to reach the vulnerable state and determine the precise concrete values that actually trigger the bug. Existing directed fuzzing approaches are ineffective at recovering the necessary trigger scaffold, while LLM-only generation is brittle because it struggles with concrete-value discovery and runtime nondeterminism. We design SyzHarness, a framework that combines LLM reasoning with coverage-guided fuzzing for patch-based Linux kernel vulnerability reproduction. Given a patch, SyzHarness uses an LLM agent grounded by code navigation tools to synthesize a parameterized fuzzing harness that fixes the prerequisite setup logic while exposing only uncertain, bug-critical input parameters to be mutated by Syzkaller. SyzHarness then translates this harness into a Syzkaller compatible interface and iteratively refines it using hierarchical reachability feedback. We evaluate SyzHarness on multiple datasets of triggerable real-world Linux kernel vulnerabilities. On 100 KernelCTF cases, SyzHarness achieves a 78% bug reproduction success rate. On the SyzDirect benchmark, SyzHarness achieves a 73% bug reproduction success rate, substantially outperforming prior directed greybox fuzzing. On 50 recent, known-triggerable syzbot bugs fixed after March 2026, SyzHarness reproduces 40/50 (80%) using only the fix commits as input.

View source

Similar papers

Preprint Sep 2026

Robustness-Aware Evaluation and Enhancement of Mutation-Based Fuzzing for Bug Discovery

Splitting, a black-box wrapper that copies a fuzzer's queue state after a bug trigger and continues from that state in multiple branches, directing more effort toward the discovered region, provides a practical way to measure and improve fuzzers.

Zi-Rui Liu, Meng-Fan Xu, Juan Zhai et al. · 0 citations
Open access Oct 2026

SmartFuzz: Leveraging Large Language Models and Feature Composition to Generate High-Quality Seeds for Database Fuzzing

Mutation-based fuzzing is one of the most effective techniques for uncovering bugs in Database Management Systems (DBMSs). However, its effectiveness critically depends on the quality of the initial seed queries. High-quality seeds should be syntactically and semantically valid, incorporate diverse SQL features, and en...

Li Lin, Jin-Tai Hong, Yan Zhuang et al. · 0 citations
Preprint Aug 2026

FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs

Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the model to generate a proof-of-concept input that triggers a predefined target vulnerability. However, this setup may overlook valid crashes disc...

Ze Sheng, Aleksandar Kezic, Zhicheng Chen et al. · 0 citations
Preprint Sep 2026

AIJon: Automated Generation of Annotations for Fuzzing

Modern fuzzers use code coverage as feedback to guide their exploration which has proven to be an effective strategy for driving exploration. However, this strategy overlooks inputs that may be interesting to the target program even without uncovering new code paths. Fortunately, prior research has shown that annotatio...

J. Vadayath, Hu-Lin Wang, Moritz Schloegel et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction

This work recasts vulnerability discovery as an input-prediction task with a closed, deterministic ground truth, and decomposes discovery into three task modes over 22 real-world C/C++ programs spanning 15 domains, finding constraint inference, not navigation, is the dominant bottleneck.

Yuan-Xiang Shi, Jia-Yi Lin, Xuan-Yong Lin et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.