CXL Memory for LLM Inference Staging via an On-Device DMA Controller
Abstract
LLM inference frequently exhausts GPU memory, forcing frameworks to stage data such as KV caches and intermediate tensors out of GPU HBM. Today this is done with GPUDirect Storage (GDS) over PCIe to NVMe SSDs, but NVMe bandwidth and latency remain a bottleneck for the fine-grained, high-frequency accesses these workloads generate. We study a CXL memory expander with an on-device DMA controller as a faster staging tier. Although current GPUs do not support CXL, our design addresses this by using the device-side DMA controller to drive PCIe peer-to-peer (P2P) transfers directly against GPU HBM, while CXL is used only on the host side for capacity expansion and management. Using a custom NIXL backend plugin and NIXLBench, we characterize GPU-to-expander transfers and compare them head-to-head with GDS to four Gen5 × 4 NVMe drives on the same host. The expander reaches ∼ 51 GB/s reads (∼ 80% of the PCIe Gen5 × 16 peak) and ∼ 29 GB/s writes. It delivers substantially higher read bandwidth than NVMe-backed GDS across the full 4 KB–8 MB range, and higher write bandwidth for block sizes up to ∼ 1 MB—the range most relevant to KV cache and tensor staging in production LLM workloads. At larger block sizes, aggregated NVMe write bandwidth becomes competitive and eventually exceeds the expander. At KV-relevant 4 KB–128 KB blocks, the expander’s end-to-end latency is more than an order of magnitude lower than GDS (single-digit vs. tens of microseconds). The expander exhibits a read/write bandwidth asymmetry; characterizing and closing this gap is a direction for future work.