Skip to content
Preprint

JITterFlip: Uncovering Fault Attack Surfaces in JIT-Compiled LLM Serving

Aug 2026 · 0 citations · 68 references
Computer Science

TL;DR

JITterFlip is presented, the first BFA targeting the host-side JIT serving control plane of GPU-based LLM inference, and develops a decision-guided fault-vulnerable code analysis that enables both gibberish output generation and a correct-output sponge attack.

Abstract

LLMs are widely deployed through cloud-hosted inference services, where Just-in-Time (JIT) compilation is used to reduce recurring framework and GPU-launch overhead. JIT serving introduces a host-side control plane that selects compiled artifacts and orchestrates their execution on the GPU. Meanwhile, the shared cloud setting has motivated a growing body of bit-flip attacks (BFAs) against LLM/DNN inference. Most existing BFAs target model parameters or weights and require model-specific knowledge. A smaller body of work reduces this dependency by faulting executable code, yet still corrupts code that directly implements model computation, limiting their attack effect to inference depletion. We present JITterFlip, the first BFA targeting the host-side JIT serving control plane of GPU-based LLM inference. By faulting CPU-resident serving decisions rather than model computation, JITterFlip enables both gibberish output generation and a correct-output sponge attack. To identify exploitable targets in a large JIT compiler stack, JITterFlip develops a decision-guided fault-vulnerable code analysis. Across four text and multimodal LLM workloads, the identified vulnerable code faults exhibit cross-model transferability, produce gibberish outputs with PPL ratios of $15.45\times$ to $2.48{\times}10^{6}\times$, and demonstrate correct-output sponge attacks with latency amplification of $2.03\times$ to $181.90\times$. JITterFlip also bypasses recent BFA defenses for LLMs while retaining both attack effects. Last, we demonstrate end-to-end Rowhammer attacks across four LLMs: a single bit flip in CPU-resident branch code propagates across the CPU-GPU boundary to disrupt GPU-executed inference without direct access to GPU memory, reaching up to $7.23{\times}10^{6}\times$ PPL amplification or $124.97\times$ latency amplification while preserving the exact generated output.

View source

Similar papers

Book Open access Sep 2026

LMTracer: Fine-Grained and Real-Time Performance Profiling for Production LLM Systems

Training and serving large language models (LLMs) has become a core business for AI providers. To ensure a high-quality user experience while optimizing infrastructure costs, providers need to closely monitor the performance of LLM executions in production. However, existing performance profiling tools fall short in th...

Wei Liu, Yong-Chao He, Bo-Han Zhao et al. · 0 citations
Book Open access Sep 2026

Quantifying the Code-Size Overhead of eBPF JIT Compilation

eBPF allows user-defined programs to safely extend Linux kernel functionality at runtime, but its final machine code comes from a compilation pipeline that differs from native targets, and how efficient that pipeline is has no clear reference point. Our work constructs one: using the standard LLVM x86 backend as an app...

Hoang Duong, Hao Sun, Zhen-Dong Su · 0 citations
Preprint Aug 2026

Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure

Design guidance for measurement under strategic optimization is distill design guidance for measurement under strategic optimization: held-out probes retain validity only on non-enumerable axes; gates must measure held-out performance, not just correctness; and a transfer rate is interpretable only with per-failure mec...

Víctor Gallego · 0 citations
Preprint Aug 2026

GPU Offload in Rust: Portable, Safe, and Fast

This paper presents a zero-overhead, multi-vendor GPU compilation framework built natively into the Rust compiler (rustc) and LLVM backends, and leverages Rust's rich type system, ownership system, and strict aliasing guarantees to efficiently manage and optimize data transfers through LLVM's Offload infrastructure.

Manuel S. Drehwald, Marcelo Domínguez, Kevin Sala et al. · 1 citation
Aug 2026

PMDangNull: Preventing use-After-Free for Persistent Memory Applications

Use-After-Free (UAF) remains one of the most critical security threats affecting C/C++ programs. Moreover, the cross-restart persistence semantics of persistent memory (PM) programming models significantly broaden the UAF attack surface. Existing DRAM-based protection schemes lack crash consistency guarantees, whereas...

Yuquan Chi, Yinjin Fu, Yong-Gang Hu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.