This article investigates two lightweight field-programmable gate array (FPGA) implementations of an iterative NTT-based polynomial multiplication accelerator through non-pipelined and 4-stage pipelined architectures, showing that the non-pipelined architecture provides reduced hardware overhead and lower power consumption, whereas the pipelined architecture improves timing scalability and successfully operates at 280 MHz.
Abstract
Number Theoretic Transform (NTT)-based polynomial multiplication is a computationally intensive operation in lattice-based post-quantum cryptography (PQC) schemes such as CRYSTALS-Dilithium. Existing hardware accelerators optimize area and timing performance, without focusing on evaluating trade-offs among hardware utilization, execution latency, operating frequency, and power consumption. This article investigates such trade-offs through two lightweight field-programmable gate array (FPGA) implementations of an iterative NTT-based polynomial multiplication accelerator, namely non-pipelined and 4-stage pipelined architectures. Both implementations employ a single butterfly unit based on Cooley–Tukey and Gentleman–Sande configurations to compute the forward NTT (FNTT), inverse NTT (INTT), and coefficient-wise multiplication (CWM). The 4-stage pipelined architecture employs pipeline registers in the modular multiplication and Barrett reduction datapaths to maximize the operating frequency. Both architectures are implemented on an Artix-7 FPGA and evaluated across operating frequencies ranging from 10 MHz to 280 MHz. The results show that the non-pipelined architecture provides reduced hardware overhead and lower power consumption, whereas the pipelined architecture improves timing scalability and successfully operates at 280 MHz. At the maximum operating frequency, the pipelined implementation utilizes 1115 slices and achieves execution times of 4.58 μs, 0.93μs, and 4.58μs for FNTT, CWM, and INTT computations, respectively, with an average power consumption of 133 mW. The Area–Time Product (ATP) and Energy–Delay Product (EDP) evaluations demonstrate that the pipelined architecture achieves improved overall efficiency within the proposed lightweight single-butterfly-based polynomial multiplication architecture at higher operating frequencies, obtaining an ATP of 7.74×103 Slice-μs and EDP of 1501.52 nJ-μs.
CRYSTALS-Kyber and CRYSTALS-Dilithium are representative lattice-based post-quantum cryptographic schemes, where number theoretic transform (NTT), inverse NTT (INTT), and point-wise multiplication (PWM) dominate polynomial arithmetic. Existing hardware accelerators are typically optimized for a single scheme or operati...
Yuchen Wang, Xiaoke Wang, Chaoxing You et al.· Electronics· 0 citations
The iterative forward and inverse number theoretic transform (NTT) is a key component in lattice-based post-quantum cryptography (PQC), typically implemented using Cooley-Tukey and Gentleman-Sande butterfly units. Existing iterative NTT accelerators often rely on ping-pong memory schemes and large memory blocks tied to...
Malik Imran, A. Khalid, C. Rafferty et al.· arXiv.org· 0 citations
A hierarchical multiplier architecture consisting of partial-product generation, 4–2 compressor-tree reduction, and a look-ahead adder is developed to improve multiplication throughput and reduce latency compared with conventional Wallace-tree implementations.
A systolic array-based accelerator prototype implemented on the Zynq-7000 SoC integrated directly into PyTorch, enabling dense linear algebra to be delegated to the FPGA chip while Cortex-A9 continues executing the software stack uninterrupted.
Omar Hernandez-Yañez, A. Juárez-Lora, J. Y. Montiel-Pérez et al.· Electronics· 0 citations
Arithmetic operations are fundamental to digital signal processing systems, where multipliers often decide overall performance constraints. They are key components of many high-performance systems such as Microprocessors, FIR Filters, Digital Signal Processors etc. The most common way of performing signed multiplicatio...
C.S. Chakradhar, M. Sreedhar· ITEGAM- Journal of Engineeri...· 0 citations
This research presents an optimized, reduced-clock-cycle approach for performing 128-bit floating-point calculations on Field-Programmable Gate Arrays (FPGAs) using the SRMA architecture, which not only reduces execution time but also decreases energy consumption on Kintex-7 FPGA boards.
N. P, Thirumalaiswamy V., H. S. et al.· International Journal of Com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.