Skip to content
Conference

A High-Speed and Low-Complexity NTT Hardware Architecture for CRYSTALS-Kyber

Aug 2026 · International Test Conference in Asia · pp. 1-6 · 0 citations · 22 references

Abstract

With the ongoing standardization of Post-Quantum Cryptography (PQC), polynomial multiplication in lattice-based cryptographic schemes has emerged as a critical performance bottleneck. The Number Theoretic Transform (NTT), as the core technique for accelerating such computations, plays a decisive role in determining the efficiency of practical PQC deployments. To address this challenge, this paper proposes a low-complexity and high-throughput hardware architecture for NTT/INTT acceleration. First, a highly parallel structure consisting of eight configurable butterfly units is designed, enabling resource sharing between NTT/INTT operations and supporting point-wise multiplication (PWM), thereby significantly improving throughput and computational efficiency. Second, an optimized Barrett reduction scheme is introduced, featuring a three-stage modular reduction unit to effectively reduce the latency of modular operations. Finally, a novel polynomial coefficient reordering strategy is proposed to optimize memory access patterns, simplify addressing logic, and reduce the complexity of the control unit. Based on these optimizations, the proposed accelerator significantly enhances the performance of the CRYSTALS-Kyber algorithm. Experimental results on an Artix-7 FPGA demonstrate that, at a clock frequency of 150 MHz, a single NTT or INTT operation requires only 248 clock cycles, while the PWM stage requires 76 cycles. A complete polynomial multiplication can be finished within 5.4μs, achieving excellent performance and efficiency.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.