Area Efficiency in Constant-Cycle Schemes
Abstract
Boolean masking is a widely used countermeasure to protect hardware implementations against side-channel attacks. However, combining security with low latency remains challenging, since in many composed masked designs the latency increases with the number of cascaded non-linear operations. Constant-cycle masking schemes address this limitation by maintaining a fixed evaluation latency for a given security order d, independent of the number of non-linear functions. In this work, we propose CCHPC1.1, an arbitrary-order masked gadget construction that improves Constant-Cycle Hardware Private Circuits (CCHPC) while preserving security in the Robust but Relaxed (RR) d-probing model and under the Probe- Isolating Non-Interference (PINI) composability notion. Building on these gadget-level improvements, we further propose architecture-level optimizations that reduce the overall costs of Dual-Rail Pre-charge (DRP) logic in constant-cycle schemes, namely LUT-based Masked Dual-Rail with Pre-charge Logic (LMDPL) and CCHPC. We integrate these optimizations into a low-latency Advanced Encryption Standard (AES) core with 10+d clock cycles of total latency by extending the existing duality concept into duality+, enabling consecutive evaluation cycles at reduced area cost. Overall, our optimized design reduces the area of this low-latency cipher core by at least 44.98% for first-order security, with increasing area savings for higher security orders.