Aug 2026· Electronics· Vol 15, pp. 3495· 0 citations· 25 references
TL;DR
Results demonstrate that the proposed architecture provides an efficient, scalable, and highly reusable hardware solution for lattice-based post-quantum cryptographic accelerators.
Abstract
CRYSTALS-Kyber and CRYSTALS-Dilithium are representative lattice-based post-quantum cryptographic schemes, where number theoretic transform (NTT), inverse NTT (INTT), and point-wise multiplication (PWM) dominate polynomial arithmetic. Existing hardware accelerators are typically optimized for a single scheme or operation, resulting in limited scalability and redundant hardware resources. This paper presents a unified and scalable NTT/INTT/PWM architecture for both Kyber and Dilithium. A configurable two-dimensional processing-element (PE) array combined with an EVEN/ODD ping-pong memory organization enables conflict-free memory access and partial inter-stage pipelining. A unified K-RED-based modular multiplier supports either one Dilithium multiplication or two parallel Kyber multiplications using the same DSP resources, while PWM is mapped onto the existing PE array without requiring a separate complete PWM arithmetic array. Experimental results on a Xilinx Artix-7 FPGA show that the proposed modular multiplier reduces LUT utilization by 39.8% compared with the previous unified Kyber/Dilithium design. Under matched PE configurations, the proposed architecture achieves ATP reductions of up to 79.5% for Dilithium NTT/INTT and 73.9% for Kyber NTT/INTT, while the Kyber/Dilithium PWM mode achieves an ATP reduction of up to 78.3%. These results demonstrate that the proposed architecture provides an efficient, scalable, and highly reusable hardware solution for lattice-based post-quantum cryptographic accelerators.
With the ongoing standardization of Post-Quantum Cryptography (PQC), polynomial multiplication in lattice-based cryptographic schemes has emerged as a critical performance bottleneck. The Number Theoretic Transform (NTT), as the core technique for accelerating such computations, plays a decisive role in determining the...
Qing-Xin Xie, Li-Ping Wang, Xiao Liu et al.· International Test Conferenc...· 0 citations
The widespread deployment of deep neural networks on edge devices faces a severe imbalance between computational demand and available power, while devices frequently switch between low-power standby and highperformance detection modes. Existing general-purpose processors, graphics processing units, and fixed-precision...
Zi-Qing Mai, Zhan-Peng Jiang· 2026 2nd International Confe...· 0 citations
The transition toward post-quantum cryptography requires hardware architectures that can execute lattice-based polynomial arithmetic with low latency and moderate resource cost. This paper presents a Verilog-based realization and evaluation of the Karatsuba Initiated Novel Accelerator (KINA) concept for Ring-Binary Lea...
Arukonda Suresh and K Naresh· International Journal of Adv...· 0 citations
The proposed mapping framework achieves reduced latency and improved area–latency trade-offs compared to prior MAGIC designs, and comparative evaluation shows competitive performance for adders and substantially lower latency with improved scalability for multiplier architectures compared with representative MAC-, MAJ-...
S. Nabipour, F. Shirinzadeh, Kamalika Datta et al.· Journal of Signal Processing...· 0 citations
This work explores the capability of 8T-2MTJ bitcell, originally investigated for ternary search operation, to support both associative search and In-Memory Computing (IMC) within the same Non-Volatile Ternary Content-Addressable Memory (NV-TCAM) framework. The proposed approach reuses the existing search-line control...
Deepak Joshi, Sukhen Mondal, Mohit Gupta et al.· International Symposium on V...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.