CopyCat: Harvesting the Frequency Tax of Bulk Memory Copy
Abstract
Bulk memory copying (memcpy) is a dominant operation in modern data centers, driven by storage engines, in-memory databases, and caches. The emergence of heterogeneous memory (CXL, persistent memory, remote NUMA) further increases the volume of concurrent memory copy operations, making it easier to saturate the shared memory bandwidth. Past saturation, a higher CPU frequency no longer improves copy throughput. However, utilization-based OS control and Intel HWP both keep copying cores at the maximum frequency in the saturated copy phases that we evaluate, wasting up to 24% of CPU package energy on memory access stall cycles. We present CopyCat, a user-space runtime with a memcpy-like submission interface and an explicit completion operation that reclaims this wasted energy. It provides router selects between two complementary paths: a synchronous mode that down-clocks copying cores in place, and an asynchronous mode that offloads copies to dedicated low-frequency cores, freeing the issuing cores for compute. In a blob-cache server prototype, CopyCat saves 8% package energy in the asynchronous serve phase due to computing-memory accessing overlap that reduce the end-to-end execution time, while dedicated low-frequency cores further saves up to 22% under 1% throughput loss from frequency reduction.