Adaptive Quantum-Classical Cryptographic Selection: An RL-Based Architecture for DN25
Abstract
We present Q-OPSEC, an adaptive middleware that uses supervised, unsupervised and reinforcement learning to select cryptographic strategies from classical, post-quantum and quantum-assisted (QKD) options. The selection is modeled as a multi-objective MDP that balances security, latency, computational and energy cost, and compliance. A negotiator and registry enforce hard constraints, handle endpoint compatibility and fallbacks, and store empirical cost profiles. Experiments in simulated smart environments and hardware-in-the-loop tests show high success rates ($>95 \%$) and context-aware adaptation. Limitations include reliance on simulated QKD channels, limited device profiling, empirically tuned hyperparameters, and evaluation in high-performance environments, which may not reflect IoT constraints; future work targets real QKD integration, broader benchmarking, robust RL methods, federated learning and explainability.