Robust Deep Reinforcement Learning via Adversarial Data Augmentation for Robotic Control Under Observation Uncertainty
Abstract
Observation uncertainty can drive a deep reinforcement learning (DRL) controller to select inappropriate actions even when its nominal policy performs well. This issue is especially pronounced in robotic systems deployed outside controlled laboratory conditions, where sensing and communication are affected by noise, bias, and intermittent errors. We introduce robust deep reinforcement learning with adversarial data augmentation (RDRL-ADA), which improves resistance to perturbed observations without sacrificing nominal control quality. The method first augments a standard Markov decision process with an uncertainty set and a state-perturbation mapping. It then instantiates worst-case and stochastic observation-robust formulations for adversarial and random disturbances. To obtain varied training samples, RDRL-ADA alternates between a Bayesian-optimization black-box search and a gradient-based white-box search, and stores intermediate as well as final perturbations in the replay buffer. Policy learning is performed in an off-policy maximum-entropy framework subject to a Tsallis-entropy constraint. Experiments on six MuJoCo control tasks show that the resulting policies maintain stronger performance under observation perturbations and, in most cases, also improve the nominal return over the comparison methods.