Skip to content

Distill-and-Rank: Evaluating Fast Transformer Models for Keyless PQC Side-Channel Recovery

Aug 2026 · Midwest Symposium on Circuits and Systems · pp. 376-380 · 0 citations · 29 references

Abstract

Post-quantum cryptography (PQC) schemes such as the NIST-standardized ML-KEM (CRYSTALS-Kyber) are entering deployment, yet their resilience to side-channel attacks (SCAs) remains insufficiently understood. We study a realistic keyless attack on ML-KEM’s pointwise multiplication: a leakage model trained only on a disjoint known-key set recovers each target’s secret by ranking candidate coefficients without ever using its key, relying solely on the public ciphertext r and the traces. On low-noise simulated traces and real CW305 FPGA measurements, we recover the NTT-domain coefficients of $s_{1}$ and benchmark CPA, a 1-D CNN, and a compact hybrid (Reg-Net) against our Distill-and-Rank framework, which pairs PatchSCA-Lite, a lightweight patch-based Transformer (153k parameters, ~2 ms/trace), with EMA self-distillation, correlation ranking, and keyless confidence measures (margin, bootstrap stability). On low-noise traces all learned models beat CPA (25 traces to disclosure); the compact hybrid needs the fewest (12) and PatchSCA-Lite 17, with EMA distillation essential for the Transformer to converge. Under high noise and on real hardware, learned models degrade sharply: the compact hybrid gives the best efficiency-robustness tradeoff, while pure attention collapses where low SNR and scale mismatch dominate. Standardized PQC implementations remain vulnerable to leakage, and lightweight hybrid models are the more practical choice for keyless SCA.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.