Skip to content
Open access

Design and Implementation of a High-Performance RISC-V CPU IP Core for Multiple Cryptographic Algorithms

Aug 2026 · Advanced Electromagnetics · 0 citations

TL;DR

A hierarchical multiplier architecture consisting of partial-product generation, 4–2 compressor-tree reduction, and a look-ahead adder is developed to improve multiplication throughput and reduce latency compared with conventional Wallace-tree implementations.

Abstract

To satisfy the low-overhead and high-throughput requirements of embedded computing platforms executing cryptographic workloads, this study presents a high-performance RISC-V CPU IP core compatible with the standard RV32IM architecture. Instead of introducing dedicated cryptographic instructions, the proposed design enhances the execution efficiency of algorithms such as SM3 and SM4 through microarchitectural optimization of arithmetic units. A hierarchical multiplier architecture consisting of partial-product generation, 4–2 compressor-tree reduction, and a look-ahead adder is developed to improve multiplication throughput, while an enhanced radix-4 SRT divider incorporating leading-zero detection, result caching, and dynamic iteration control is proposed to reduce division latency. All modules are implemented in Verilog and integrated into a five-stage RISC-V pipeline. Functional verification is performed using a complete RVM instruction test suite, followed by FPGA-based performance evaluation. Experimental results demonstrate that the multiplier generates 64-bit results within a single clock cycle in 32-bit scenarios, reducing latency by more than 70% compared with conventional Wallace-tree implementations. The proposed architecture provides an efficient computing platform for secure embedded systems and offers implementation references for real-time signal processing, communication security, and electromagnetic information systems.

Read PDF

Similar papers

Open access Sep 2026

Design and FPGA Evaluation of a Multi-Operand Extension Mechanism for Tightly Coupled RISC-V Processors

Complex scalar kernels often contain instruction fragments whose computation is simple but whose operand demand exceeds the traditional two-input–one-output custom-instruction model. This operand interface bottleneck limits instruction merging in tightly coupled RISC-V acceleration. This paper proposes a multi-operand...

Peng Lu, Mei-Jiao Yu, You-Ping Mao et al. · 0 citations
Conference Aug 2026

High-Performance Implementation of AES-128 Encryption and Decryption on FPGA Using Pipeline Architecture with UART Communication

With the growing number of sophisticated cryptographic attacks and the need for high-speed secure data communications, the design and implementation of hardware-based encryption appears inevitable. This paper aims to design a highspeed pipelined Advanced Encryption Standard (AES-128) architecture using the latest 28 nm...

H. K, Nihal G. Deshakulkarni, Nimish Vihan · 0 citations
Conference Aug 2026

Lightweight Bit Manipulation Extension Architecture and Performance Evaluation for RISC-V Processors Based on the NICE Interface

The open and modular RISC V instruction set architecture supports flexible domain specific extensions. However, conventional embedded RISC V processors rely on general purpose instruction sequences to implement bit manipulation operations, which suffer from long execution cycles, low computational efficiency, and high...

Chuan-Rong Yang, Zi-Yang Hu, Yue-Jun Zhang · 0 citations
Open access Aug 2026

High-Performance SM2 Signature Hardware Architecture Based on Precomputation and Parallel Scheduling

In high-throughput, high-concurrency, low-latency cryptographic scenarios, the throughput of public-key cryptography is a decisive performance metric. As China’s national elliptic curve cryptography standard, the SM2 signature algorithm has been widely adopted, yet scalar multiplication—the core primitive of SM2 signat...

Jie Huang, Ming-Fu Zhong, Zuo-Nan Xiao · 0 citations
Conference Aug 2026

A Highly Flexible and Reconfigurable Instruction-Driven Hardware Acceleration System for ML-DSA

The rapid advancement of quantum computing severely threatens conventional public-key infrastructure, promoting the migration to Post-Quantum Cryptography (PQC). With NIST standardizing Module-Lattice-Based Digital Signature Algorithm (ML-DSA) as FIPS 204, the high computational complexity of lattice-based cryptography...

Xiao Liu, Qing-Xin Xie, Hui-Hong Zhang et al. · 0 citations
Conference Aug 2026

Design and ASIC Implementation of a Reconfigurable Hybrid Approximate MAC Architecture for Energy-Efficient Computing

This work presents a suitable design as well as implementing a reconfigurable 32-bit Multiply-Accumulate (MAC) unit based on a hybrid approximate adder architecture aimed at energy-efficient VLSI applications. The proposed design partitions the adder into three segments according to bit significance: an 8-bit approxima...

Tangella Haswanth, Debashish Dash · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.