Reinforcement Learning for Quantum Error Correction
This work investigates the use of reinforcement learning to perform QEC on a rotated surface code as a partially observable Markov decision process (POMDP), and finds that the tabular agent is limited by the exponential growth of the history-including state-action space, whereas PPO is able to generalize across similar...