Skip to content

OptiPrime: Optimizing Private Inference through Protocol-Hardware Co-design

Sep 2026 · 0 citations · 82 references
Computer Science

TL;DR

OptiPrime is introduced, a protocol-hardware co-optimization framework for efficient private DNN inference that features a novel HE protocol for convolutions that substantially reduces the number of transmitted output ciphertexts and mitigates the network communication bottleneck.

Abstract

Private deep neural network (DNN) inference based on hybrid homomorphic encryption (HE) and multi-party computation (MPC) can protect user data with a formal guarantee, but at the cost of significant latency overhead due to HE. Customized HE accelerators have been proposed and have achieved orders-of-magnitude speedup for individual HE operations. However, when directly applying a commercial HE accelerator to state-of-the-art HE-MPC frameworks, we observe only limited end-to-end performance gain. This is because HE-MPC frameworks often require wireless transmission of input and output ciphertexts for each HE operation, leading to a severe network communication bottleneck. To overcome this challenge, we introduce OptiPrime, a protocol-hardware co-optimization framework for efficient private DNN inference. OptiPrime features a novel HE protocol for convolutions that substantially reduces the number of transmitted output ciphertexts and mitigates the network communication bottleneck. Meanwhile, as the new protocol introduces complex computation for fewer output ciphertext, we observe new memory access challenges due to a high volume of weight plaintexts and intermediate ciphertexts. Hence, we further propose a lightweight compression system for the weight plaintexts, reducing memory traffic by 10 times, as well as a specialized dataflow to maximize on-chip data reuse of intermediate ciphertexts. Extensive experiments show that our framework outperforms the Cheetah baseline by at most 5.7 times on CPUs and 4.2 times with an accelerator.

View source

Similar papers

Open access Aug 2026

EHEIR: Efficient Homomorphic Encrypted Inference via Architectural Redesign

This work presents a framework that reformulates HE-aware model design as a constrained neural architecture search problem, where the objective is to identify architectures that are both cryptographically feasible and computationally efficient while preserving task performance.

Reeshav Chowdhury, Anoop Mishra, Deepak Khazanchi et al. · 0 citations
Aug 2026

S2MM: Scalable FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption

S2MM is presented, the first scalable FPGA-based accelerator designed for HE MM, and a novel datapath for Homomorphic Linear Transformation (HLT), the dominant workload in HE MM is proposed, enabling fine-grained on-chip data reuse and substantially reducing both off-chip memory traffic and on-chip memory demand.

Zhi-Han Xu, Rajgopal Kannan, Viktor K. Prasanna · 0 citations
Book Open access Sep 2026

Cerium: A Multi-GPU Framework for Terabyte-Scale Encrypted Inference

Encrypted AI using fully homomorphic encryption (FHE) enables inference directly over encrypted queries, providing strong privacy guarantees. However, its computational and memory overheads have limited practical deployment. Custom FHE accelerators improve performance, but rely on advanced manufacturing technologies th...

Siddharth Jayashankar, Joshua Kim, Michael B. Sullivan et al. · 0 citations
Oct 2026

WestLake: Accelerating Fully Homomorphic Encryption With Less On-Chip Memory

Fully homomorphic encryption (FHE) enables computation on encrypted data without decryption. This makes FHE a valuable privacy-preserving technique applicable in fields such as private machine learning (ML). FHE achieved unlimited homomorphic operations on ciphertext by periodic bootstrapping, which is highly time-cons...

Peng-Cheng Qiu, Bao-Ze Zhao, Gui-Ming Wu et al. · 0 citations
Preprint Sep 2026

Memory-Efficient Designs for Word-Wise Universal Fully Homomorphic Encryption

Fully Homomorphic Encryption (FHE) enables computation on encrypted data, preserving privacy throughout analysis. While its privacy is very strong, FHE is much slower to execute than the original computation. In particular, due to the recent success in accelerating its compute, the performance bottleneck shifts to the...

A. W. B. Yudha, Erwin Eko Wahyudi, R. Rajagede et al. · 0 citations
Preprint Sep 2026

FPGA Acceleration of Fully Homomorphic Encryption with Adaptive Key Switching

Fully Homomorphic Encryption (FHE) enables privacy-preserving cloud services but incurs substantial computation overhead, making hardware acceleration essential. Among FHE operations, key-switching is a major performance bottleneck. Recent cryptographic advances introduce a novel key-switching method (i.e., KLSS) that...

Zhi-Han Xu, Jayashree Adivarahan, Rajgopal Kannan et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.