Sep 2026· IACR Transactions on Cryptographic Hardware and Embedded Systems· Vol 2026, pp. 970-997· 0 citations· 48 references
TL;DR
This work proposes a unified framework to efficiently support Number Theoretic Transform (NTT) workloads in both FHE and ZKP, and designs limb-wise 256-bit arithmetic to efficiently support ZKP workloads using RVV.
Abstract
Data privacy has become increasingly critical in modern society, driving significant interest from both academia and industry in privacy-preserving technologies such as fully homomorphic encryption (FHE) and zero-knowledge proofs (ZKP). More recently, emerging paradigms such as verifiable FHE require the joint support of both cryptographic techniques, significantly increasing the diversity and complexity of underlying computational workloads. Prior hardware accelerators mainly target a single application or a narrow parameter range, making them ill-suited for supporting multiple cryptographic primitives. As an emerging open platform, RISC-V combines general-purpose programmability with high-performance computation enabled by vector extensions such as RISC-V Vector Extension (RVV). In this work, we propose a unified framework to efficiently support Number Theoretic Transform (NTT) workloads in both FHE and ZKP. We observe that existing RVV-based NTT kernels fail to scale efficiently to large parameter sizes; while the four-step NTT algorithm reduces the transform size through decomposition, conventional parameter selection often leads to suboptimal performance. To address this, we introduce a model-guided decomposition strategy that automatically selects near-optimal parameters. In addition, we design limb-wise 256-bit arithmetic to efficiently support ZKP workloads using RVV. We implement our framework on a gem5-based out-of-order RISC-V core with RVV v1.0 support. Experimental results demonstrate that our approach significantly improves NTT performance across a wide range of parameters and provides an efficient execution substrate for both FHE and ZKP workloads. Specifically, our approach achieves up to 1.53x speedup over prior state-of-the-art RVV-based implementations on standalone NTTs, up to 6.15x speedup over OpenFHE on CKKS bootstrapping, and up to 1.35x speedup over libsnark on BN128-based workloads.
Encrypted AI using fully homomorphic encryption (FHE) enables inference directly over encrypted queries, providing strong privacy guarantees. However, its computational and memory overheads have limited practical deployment. Custom FHE accelerators improve performance, but rely on advanced manufacturing technologies th...
Siddharth Jayashankar, Joshua Kim, Michael B. Sullivan et al.· Proceedings of the ACM SIGOP...· 0 citations
With the growing adoption of fully homomorphic encryption (FHE), improving its computational efficiency has become a key research focus, and designing dedicated accelerators remains an effective approach. Existing FHE accelerators achieve high performance but are often constrained by empirical design tendencies that fa...
Ying-Hao Yang, Fu-Yao Liu, Jin-Kai Zhang et al.· International Test Conferenc...· 0 citations
Trusted Execution Environments (TEEs) have been extensively adopted to facilitate various security-critical applications. However, in stark contrast to well-established TEE technologies such as Intel SGX, AMD SEV, and ARM TrustZone, the adoption within the emerging RISC-V architecture has seen limited traction. Specifi...
Yu Zhao, Jia-Bei He, Ming-Ru Xu et al.· ACM Transactions on Architec...· 0 citations
Although Fully Homomorphic Encryption (FHE) enables computation over encrypted data, its substantial computational and storage overhead remains a major obstacle to practical deployment. Among available hardware platforms, FPGAs offer a favorable balance of performance, flexibility, and energy efficiency, making them a...
Lingyu Gong, Farhad Merchant· ACM Transactions on Reconfig...· 0 citations
This work presents a hardware-oriented design for accelerating the Monolith hash function on FPGA and proposes a dual-architecture framework consisting of a serial architecture and a parallel architecture to address different performance and resource constraints.
Cheng Chen, Gang-Qiang Yang, Hong-Chao Zhou et al.· ACM Transactions on Reconfig...· 0 citations