Virtually Contiguous Host Allocation for Fragmentation-Resilient CPU–GPU Transfers: Performance and Energy Analysis
Abstract
In NVIDIA Compute Unified Device Architecture (CUDA) applications, host–device data transfers are frequently orchestrated over multiple discontiguous host buffers, especially when data structures are built through numerous dynamic allocations. Even when the total transferred size is fixed, fragmentation multiplies data copy and memory pinning operations on the CPU side, becoming a first-order overhead in both time and energy consumption. This paper investigates virtually contiguous host allocation as a practical mechanism to mitigate this orchestration cost. We evaluate VCMalloc, a contiguity-preserving allocator that maintains virtual contiguity across allocation, reallocation, and free operations, allowing the logical host dataset to be transferred and registered as a single contiguous region. Experiments are conducted on an NVIDIA RTX 3060 GPU, transferring a fixed 12 GB dataset across fragmentation levels ranging from 1 to over 2 million fragments. The results show that conventional allocators (Malloc, Mimalloc, cudaHostAlloc) degrade sharply as fragment count increases and collapse beyond 64K fragments — at which point memory registration alone exceeds 78 seconds. In contrast, VCMalloc maintains registration overhead below 80 milliseconds and keeps transfer time and energy near optimal, independently of the fragmentation level.