Safe Offline Reinforcement Learning for Autonomous Driving via Causal Risk Features
Abstract
Offline safe reinforcement learning aims to learn constraint-satisfying policies from pre-collected datasets without online interaction, which is critical for safety-critical applications such as autonomous driving. However, offline datasets collected from heterogeneous sources often contain spurious correlations between cost-irrelevant state features and safety signals, which can mislead policy learning and cause constraint violations under distribution shift. To address this issue, we propose a two-stage framework: (1) a causal model is employed to identify state features that are necessary and sufficient for safety constraints; (2) a safe policy is trained on the extracted features using Hamilton-Jacobi reachability-based cost value estimation and gradient-based action correction. Experiments on the MetaDrive autonomous driving benchmark demonstrate that our method achieves state-of-the-art performance in both reward and cost constraint satisfaction, and ablation study confirms that causal feature extraction significantly improves policy performance and training stability in complex driving scenarios. This work highlights the promise of causal representation learning as a principled approach to improving both performance and safety in offline reinforcement learning.