Jul 2026· IACR Transactions on Cryptographic Hardware and Embedded Systems· Vol 2026, pp. 198-223· 0 citations· 28 references
TL;DR
BOLT-FHE shows that a portable, fused-kernel organization with explicit on-chip budgeting can substantially improve TFHE bootstrapping throughput while remaining compatible with both noise management paths.
Abstract
Bootstrapping is the main performance bottleneck in bitwise Fully Homomorphic Encryption (FHE), and practical acceleration requires careful orchestration of the blind rotation and external product chain under GPU resource constraints. This paper presents BOLT-FHE, a GPU bootstrapping framework that emphasizes block-local execution, on-chip tiling, and a unified MegaKernel supporting both gadget decomposition and modulus raising, with optional support for a recently proposed technique (Bergerat et al., CHES 2025) based on the common mask assumption (CM packing). Our design keeps the accumulator update chain within a single thread block and fuses NTT/INTT, external products, and accumulator updates using a fixed execution template. Two compile-time parameters—WPP (warps per polynomial) and IPT (items per thread)—control multi-warp cooperation and per-thread register footprint, enabling consistent kernel structure across different parameter sets.On an NVIDIA RTX 4090, BOLT-FHE reaches 40,166 bootstrappings per second at 128-bit security, demonstrating high-throughput TFHE bootstrapping on a commodity GPU. Compared to the state-of-the-art GPU implementation VeloFHE (Shen et al., CHES 2025), BOLT-FHE achieves 1.01x–2.92x speedups with gadget decomposition. In particular, for modulus raising, BOLT-FHE improves by 2.38x–2.42x without CM packing, and by 3.17x–3.31x under the best packing configuration, reflecting the combined benefits of fused arithmetic, more regular memory access, and amortization enabled by CM packing. Overall, BOLT-FHE shows that a portable, fused-kernel organization with explicit on-chip budgeting can substantially improve TFHE bootstrapping throughput while remaining compatible with both noise management paths.
The iterative forward and inverse number theoretic transform (NTT) is a key component in lattice-based post-quantum cryptography (PQC), typically implemented using Cooley-Tukey and Gentleman-Sande butterfly units. Existing iterative NTT accelerators often rely on ping-pong memory schemes and large memory blocks tied to...
Malik Imran, A. Khalid, C. Rafferty et al.· arXiv.org· 0 citations
This work presents a hardware–software co-design whose 16-bit tag uses odd parity and fail-safe class encoding, which combines a variable-precision extent, an aligned CRC checker, and an exact-bounds micro-cache.
Dan Toderici, T. Enache, R. Rughinis et al.· Computers· 0 citations
Hardware for lattice-based post-quantum cryptography spends a large share of its area on the number-theoretic transform (NTT), dominated by modular multipliers and twiddle storage. FoldNTT is a redesign of the released radix-2 CFNTT accelerator (TCHES 2022) for the Falcon / FN-DSA prime q = 12289 with one hardware mult...
Although Fully Homomorphic Encryption (FHE) enables computation over encrypted data, its substantial computational and storage overhead remains a major obstacle to practical deployment. Among available hardware platforms, FPGAs offer a favorable balance of performance, flexibility, and energy efficiency, making them a...
Lingyu Gong, Farhad Merchant· ACM Transactions on Reconfig...· 0 citations
This paper describes DeepSig's CUDA-based acceleration backend for the OCUDU physical layer and O-RAN fronthaul path, integrated through acceleration interfaces that are largely independent of the underlying acceleration mechanism. The design accelerates PDSCH, PUSCH, SRS, PRACH, split-8 lower-PHY transforms, and O-RAN...
M. Pennybacker, Wanze Liu, A. Kharchenko et al.· 2 citations
This paper presents a high-performance SIMD acceleration framework for the Steinhaus-Johnson-Trotter algorithm, targeted at modern x86-64 architectures using the AVX2 instruction set. By exploiting a novel combinatorial space partitioning combined with single-cycle vector byte shuffling ($\texttt{\_mm256\_shuffle\_epi8...
S. Mel'nikov· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.