OptiPrime is introduced, a protocol-hardware co-optimization framework for efficient private DNN inference that features a novel HE protocol for convolutions that substantially reduces the number of transmitted output ciphertexts and mitigates the network communication bottleneck.
Abstract
Private deep neural network (DNN) inference based on hybrid homomorphic encryption (HE) and multi-party computation (MPC) can protect user data with a formal guarantee, but at the cost of significant latency overhead due to HE. Customized HE accelerators have been proposed and have achieved orders-of-magnitude speedup for individual HE operations. However, when directly applying a commercial HE accelerator to state-of-the-art HE-MPC frameworks, we observe only limited end-to-end performance gain. This is because HE-MPC frameworks often require wireless transmission of input and output ciphertexts for each HE operation, leading to a severe network communication bottleneck. To overcome this challenge, we introduce OptiPrime, a protocol-hardware co-optimization framework for efficient private DNN inference. OptiPrime features a novel HE protocol for convolutions that substantially reduces the number of transmitted output ciphertexts and mitigates the network communication bottleneck. Meanwhile, as the new protocol introduces complex computation for fewer output ciphertext, we observe new memory access challenges due to a high volume of weight plaintexts and intermediate ciphertexts. Hence, we further propose a lightweight compression system for the weight plaintexts, reducing memory traffic by 10 times, as well as a specialized dataflow to maximize on-chip data reuse of intermediate ciphertexts. Extensive experiments show that our framework outperforms the Cheetah baseline by at most 5.7 times on CPUs and 4.2 times with an accelerator.
This work presents a framework that reformulates HE-aware model design as a constrained neural architecture search problem, where the objective is to identify architectures that are both cryptographically feasible and computationally efficient while preserving task performance.
Reeshav Chowdhury, Anoop Mishra, Deepak Khazanchi et al.· ACM Transactions on Internet...· 0 citations
S2MM is presented, the first scalable FPGA-based accelerator designed for HE MM, and a novel datapath for Homomorphic Linear Transformation (HLT), the dominant workload in HE MM is proposed, enabling fine-grained on-chip data reuse and substantially reducing both off-chip memory traffic and on-chip memory demand.
Zhi-Han Xu, Rajgopal Kannan, Viktor K. Prasanna· ACM Transactions on Reconfig...· 0 citations
Encrypted AI using fully homomorphic encryption (FHE) enables inference directly over encrypted queries, providing strong privacy guarantees. However, its computational and memory overheads have limited practical deployment. Custom FHE accelerators improve performance, but rely on advanced manufacturing technologies th...
Siddharth Jayashankar, Joshua Kim, Michael B. Sullivan et al.· Proceedings of the ACM SIGOP...· 0 citations
Fully homomorphic encryption (FHE) enables computation on encrypted data without decryption. This makes FHE a valuable privacy-preserving technique applicable in fields such as private machine learning (ML). FHE achieved unlimited homomorphic operations on ciphertext by periodic bootstrapping, which is highly time-cons...
Peng-Cheng Qiu, Bao-Ze Zhao, Gui-Ming Wu et al.· IEEE Transactions on Very La...· 0 citations
Fully Homomorphic Encryption (FHE) enables computation on encrypted data, preserving privacy throughout analysis. While its privacy is very strong, FHE is much slower to execute than the original computation. In particular, due to the recent success in accelerating its compute, the performance bottleneck shifts to the...
A. W. B. Yudha, Erwin Eko Wahyudi, R. Rajagede et al.· 0 citations
Exploring how generative AI could make machine vision more accessible to businesses. The post GenEye in a Box: Making Machine Vision Something You Can Just Ask For appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduOct 7, 2026
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026