Optimized Design of a Scalable Post-Quantum Cryptography Accelerator Based on 2PtNTT
Abstract
To address high latency and parameter heterogeneity in post-quantum cryptography, this paper proposes a scalable hardware acceleration scheme. To mitigate the polynomial multiplication bottleneck, we introduce the 2 Preprocess-then-NTT (Number Theoretic Transform) architecture. By decomposing complex polynomial transformations into medium-scale parallel tasks and configuring multiple butterfly units to execute them concurrently, this architecture outperforms the serial computation limitations of traditional NTTs, reducing the NTT operation cycle time by 2/3. Furthermore, the K2-RED reduction algorithm is optimized for diverse lattice-based parameter sets, and the Secure Hash Algorithm 3 logic is unified to support multi-mode hashing. Experimental results on the FPGA (Field-Programmable Gate Array) demonstrate a 57.9% overall reduction in computational latency. This architecture achieves an exceptional trade-off between throughput, hardware overhead, and algorithmic flexibility, providing a robust foundation for high-performance Post-Quantum Cryptography chips.