Skip to content
Open access

High-Performance SM2 Signature Hardware Architecture Based on Precomputation and Parallel Scheduling

Aug 2026 · Electronics · 0 citations · 10 references

Abstract

In high-throughput, high-concurrency, low-latency cryptographic scenarios, the throughput of public-key cryptography is a decisive performance metric. As China’s national elliptic curve cryptography standard, the SM2 signature algorithm has been widely adopted, yet scalar multiplication—the core primitive of SM2 signature—constitutes the dominant latency bottleneck. This paper proposes a high-throughput ASIC architecture that integrates precomputation with parallel task scheduling to accelerate SM2 signature generation. First, we design a Comb-algorithm-based precomputation hardware scheme for fixed-base scalar multiplication. The 256-bit scalar is partitioned into 32 segments, and 32 dedicated lookup tables are precomputed in on-chip SRAMs, which reduces fixed-base scalar multiplication to at most 31 elliptic curve point additions. Second, a six-arithmetic-unit parallel scheduling framework is developed for variable-base scalar multiplication. Equipped with two three-stage pipelined Montgomery multipliers and four modular adders, the design overlaps point addition and point doubling across pipeline stages to boost hardware resource utilization. Moreover, we build a 16-core parallel computing platform integrated with hardware task queues and DMA automatic scheduling, achieving efficient throughput scalability with the increase in core count. The proposed design has been taped out in the TSMC 28 nm CMOS process with completed physical design, including placement, clock tree synthesis, and routing. Post-layout simulation results, with full parasitic extraction (RC) and static timing analysis (STA), demonstrate that a single core achieves 89,593 signatures per second at a post-layout maximum frequency of 600 MHz. Its normalized area efficiency, measured as signatures per kilo gate equivalent (KGE), reaches 33.31 sig/s/KGE under the TSMC 28 nm process at 600 MHz. It should be noted that this metric is significantly influenced by the advanced process node and higher operating frequency; therefore, to enable a fairer assessment of intrinsic microarchitectural efficiency independent of process scaling, the frequency-normalized metric (sig/s/MHz/KGE) is adopted as the primary cross-design benchmark. Under this metric, our design achieves 55.5 × 10−3 sig/s/MHz/KGE, which is comparable to the 55.0 × 10−3 sig/s/MHz/KGE of the most area-efficient referenced design, with a marginal improvement of approximately 0.9%. The area efficiency comparison is presented only as a supplementary reference within a limited and clearly defined scope, acknowledging that the compared designs differ in functionality, process technology, and evaluation methodology. Furthermore, our design achieves the lowest Area–Time (AT) product of 30.03 KGE·ms among the compared works, indicating that our architectural innovation achieves a favorable AT trade-off for high-throughput applications rather than a fundamental shift in circuit efficiency. The 16-core parallel computing platform reaches an overall throughput of 1.03 million signatures per second. In addition, first-order arithmetic masking and key blinding are integrated into the scalar-multiplication data path to resist first-order side-channel attacks, and simulation-based TVLA evaluation indicates its leakage suppression capability under simulated conditions.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.