Skip to content
Book Open access

Cerium: A Multi-GPU Framework for Terabyte-Scale Encrypted Inference

Sep 2026 · Proceedings of the ACM SIGOPS 32nd Symposium on Operating Systems Principles · 0 citations · 45 references

Abstract

Encrypted AI using fully homomorphic encryption (FHE) enables inference directly over encrypted queries, providing strong privacy guarantees. However, its computational and memory overheads have limited practical deployment. Custom FHE accelerators improve performance, but rely on advanced manufacturing technologies that limit their accessibility. GPUs offer a more widely available alternative, yet achieving ASIC-class performance on GPUs is challenging. Large models such as LLMs compound these challenges by requiring optimized kernels, terabyte-scale memory management, and efficient execution across multiple devices. We present Cerium, a multi-GPU framework for large-scale FHE inference. Cerium combines a domain-specific language, an optimizing compiler, and a runtime system to automatically generate optimized GPU kernels, compress plaintext data, manage terabyte-scale memory footprints, and optimize multi-GPU communication. Implemented on NVIDIA GPUs, Cerium outperforms expert-written GPU libraries by up to 2.25× and achieves performance competitive with state-of-the-art FHE ASICs, matching the CraterLake ASIC. It is the first GPU system to execute CKKS bootstrapping in under 10 milliseconds, achieving 7.5 milliseconds, and is the first to demonstrate encrypted inference for BERT-Base and Llama3-8B prefill in 8 seconds and 43 seconds, respectively. The Cerium framework is open-source and available at https://github.com/CMU-CAOS/Cerium

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.