GPU-Accelerated Post-Quantum OTA Key Exchange Using Batch-Optimized ML-KEM
Abstract
Secure over-the-air (OTA) firmware updates are indispensable to connected and autonomous vehicles (CAVs), enabling rapid vulnerability patching without physical recall. The impending arrival of cryptographically relevant quantum computers, however, threatens the public-key protocols that protect these updates. Although the National Institute of Standards and Technology (NIST) has standardized ML-KEM (FIPS 203, formerly CRYSTALS-Kyber) as the primary Post-Quantum Cryptography (PQC) Key Encapsulation Mechanism (KEM), its polynomial arithmetic exposes a critical computational bottleneck at fleet scale. Central OTA servers handling tens of thousands of simultaneous vehicle connections suffer from CPU exhaustion due to sequential Number Theoretic Transform (NTT) operations—a bottleneck that existing hardware-acceleration literature has not addressed in the context of large-scale vehicular networks. This paper proposes a high-throughput, batch-optimized GPU acceleration architecture for server-side ML-KEM-1024 key exchange while providing a realistic system-level bottleneck evaluation. By pushing isolated GPU polynomial arithmetic to a peak mathematical speedup of 92.7x (164 GOPS) at N = 32,768 clients, we definitively eliminate the lattice-math bottleneck to expose the system’s true limitation: unaccelerated CPU Keccak hashing, which our comprehensive profiling reveals dominates 95% of total execution time. Consequently, in accordance with Amdahl’s Law, the end-to-end system speedup is bounded at $1.05\times $ , yielding an overall amortized system processing time of $0.218~\mu $ s per client. PCIe data transfer overhead remains minimal at 0.027% of the total pipeline. These findings highlight that polynomial GPU offloading must be paired with Keccak acceleration to achieve true full-system throughput in automotive CSMS deployments compliant with UN R156.