A Highly Flexible and Reconfigurable Instruction-Driven Hardware Acceleration System for ML-DSA
Abstract
The rapid advancement of quantum computing severely threatens conventional public-key infrastructure, promoting the migration to Post-Quantum Cryptography (PQC). With NIST standardizing Module-Lattice-Based Digital Signature Algorithm (ML-DSA) as FIPS 204, the high computational complexity of lattice-based cryptography poses strict demands on hardware implementations. Traditional hardwired accelerators lack post-fabrication flexibility, whereas RISC-V-based co-designs suffer from excessive logic overhead and latency. This paper proposes a microcode-decoupled Application-Specific Instruction-Set Processor (ASIP) architecture. By decoupling complex control flow from data path, replacing multiplier-heavy arithmetic arrays, and adopting conflict-aware multi-bank memory scheduling, dual-core parallel processing significantly enhances efficiency and resource utilization. Implemented on Xilinx Zynq-7000 FPGA, the design operates at 125 MHz. For three security levels, it occupies 44108 LUTs, 9850 FFs, 18 BRAMs, and 0 DSPs, with total latency of 228 μs, 359 μs, and 358 μs respectively. With superior performance and resource efficiency, this work provides a high-efficiency, cost-effective hardware solution for practical PQC deployment.