Skip to content
Open access

TP-NTT: Batch NTT Hardware with Application to Relinearization

Sep 2026 · IACR Cryptology ePrint Archive · Vol 2026, pp. 556 · 0 citations · 37 references
Computer Science

TL;DR

TP-NTT is presented, a scalable, throughput-optimized NTT architecture supporting a wide range of ring dimensions used in FHE, and a relinearization accelerator is proposed that leverages the fast batch NTT capability of TP-NTT, achieving 67.34x speedup over state-of-theart software implementations and highlighting TP-NTT’s effectiveness in real-world FHE applications.

Abstract

Fully Homomorphic Encryption (FHE) enables arbitrary computation on encrypted data without decryption, providing strong privacy guarantees for secure cloud computing, encrypted analytics, and privacy-preserving machine learning. However, practical deployment of FHE remains limited by the high computational cost of polynomial arithmetic over large modular rings. In particular, Number Theoretic Transform (NTT)–based polynomial multiplication dominates the execution time of modern lattice-based FHE schemes. In this work, we present TP-NTT, a scalable, throughput-optimized NTT architecture supporting a wide range of ring dimensions used in FHE, from 210 to 216. Our design applies optimizations at multiple levels, from modular arithmetic to the NTT algorithm itself, including multi-dimensional decomposition without requiring additional multiplication blocks. The decomposition dimensionality is configurable at design time, supporting 2-D, 3-D, and 4-D decompositions, each advantageous in specific scenarios. Furthermore, TP-NTT provides design-time-configurable throughput. Combined with its scalable architecture, this enables significant advantages for batch NTT operations compared to other works in the literature. At n = 216, it outperforms the best-performing prior design by 1.33x in average latency while achieving 1.24x better area–time-product (ATP). To demonstrate its efficiency, we present a case study on FHE relinearization, focusing on the BFV scheme. We propose a relinearization accelerator that leverages the fast batch NTT capability of TP-NTT, achieving 67.34x speedup over state-of-theart software implementations and highlighting TP-NTT’s effectiveness in real-world FHE applications.

Read PDF

Similar papers

Aug 2026

S2MM: Scalable FPGA Acceleration of Secure Matrix Multiplication with Homomorphic Encryption

Homomorphic Encryption (HE) enables secure computation on encrypted data, addressing privacy concerns in cloud computing. However, the high computational cost of HE operations, particularly matrix multiplication (MM), remains a major barrier to its practical deployment. Accelerating Homomorphic Encrypted MM (HE MM) is...

Zhi-Han Xu, Rajgopal Kannan, Viktor K. Prasanna · 0 citations
Preprint Sep 2026

Memory-Efficient Designs for Word-Wise Universal Fully Homomorphic Encryption

Fully Homomorphic Encryption (FHE) enables computation on encrypted data, preserving privacy throughout analysis. While its privacy is very strong, FHE is much slower to execute than the original computation. In particular, due to the recent success in accelerating its compute, the performance bottleneck shifts to the...

A. W. B. Yudha, Erwin Eko Wahyudi, R. Rajagede et al. · 0 citations
Conference Aug 2026

Distributed Edge-Adaptive Parallel Binary Fully Homomorphic Encryption Framework

The recent growth in the privacy-sensitive artificial intelligence of distributed cloud-edge systems has accelerated the necessity of the implementation of efficient and thermally feasible encrypted inference engines. Fully Homomorphic Encryption (FHE) makes it possible to perform computation on encrypted data, and its...

Yagnasri Ashwini, S. Shailaja, K. V. N. Valli et al. · 0 citations
Preprint Sep 2026

Batched Paillier-Based Hamming-Distance Computation over Binary Embeddings

Additively homomorphic encryption supports outsourced computation on encrypted binary embeddings, but large-integer arithmetic and data movement can limit throughput. We describe a Paillier-based client that combines a carry-separated binary encoding, table-based encryption, reduced-exponent decryption, CUDA/CGBN arith...

Yavor Litchev, Li-Wen Ouyang · 0 citations
Preprint Sep 2026

FPGA Acceleration of Fully Homomorphic Encryption with Adaptive Key Switching

Fully Homomorphic Encryption (FHE) enables privacy-preserving cloud services but incurs substantial computation overhead, making hardware acceleration essential. Among FHE operations, key-switching is a major performance bottleneck. Recent cryptographic advances introduce a novel key-switching method (i.e., KLSS) that...

Zhi-Han Xu, Jayashree Adivarahan, Rajgopal Kannan et al. · 0 citations
Oct 2026

WestLake: Accelerating Fully Homomorphic Encryption With Less On-Chip Memory

Fully homomorphic encryption (FHE) enables computation on encrypted data without decryption. This makes FHE a valuable privacy-preserving technique applicable in fields such as private machine learning (ML). FHE achieved unlimited homomorphic operations on ciphertext by periodic bootstrapping, which is highly time-cons...

Peng-Cheng Qiu, Bao-Ze Zhao, Gui-Ming Wu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.