Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· 0 citations· 37 references
TL;DR
A framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator is developed, yielding a persistent safety set from which the agent can remain safe indefinitely, and a new reward maximization algorithm is proposed that effectively exploits the learned persistent safety set for reward critic estimation.
Abstract
Offline safe reinforcement learning learns high-return policies that satisfy hard safety constraints using only a pre-collected dataset. This setting is challenging due to the inability to explore, and the risk of propagating value errors through unsafe state-space regions. To address this, first, we characterize the safe state region by developing a framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator, yielding a persistent safety set, from which the agent can remain safe indefinitely. Second, we show that several existing safety set estimation methods (e.g., reachability-constrained RL) can be formulated within our CBF learning framework, highlighting its generality. We further propose a new CBF that ensures safety under environment dynamics uncertainty, unlike standard CBFs designed for deterministic settings. Third, we propose a new reward maximization algorithm that effectively exploits our learned persistent safety set for reward critic estimation. Empirical results on standard benchmarks show that our approach achieves state-of-the-art safety with fewer constraint violations while maintaining competitive returns.
A counterexample-guided reinforcement learning method that navigates safe exploration in autonomous systems without prior knowledge, even when safety and optimality conflict, and a novel belief-based regularization method to address the distributional shift between online and offline learning and to balance optimizatio...
Xiao-Tong Ji, Antonio Filieri· ACM Transactions on Autonomo...· 0 citations
Offline safe reinforcement learning aims to learn constraint-satisfying policies from pre-collected datasets without online interaction, which is critical for safety-critical applications such as autonomous driving. However, offline datasets collected from heterogeneous sources often contain spurious correlations betwe...
Zi-Qian Wang, Zhen Zhang· IEEE/ASME International Conf...· 0 citations
This paper proposes a safe meta-RL framework that explicitly accounts for safety during adaptation, and develops a safe meta-RL algorithm that learns the safety value function and leverages it for safety filtering and constrained policy optimization.
Evaluation metrics for safe RL are introduced that address each of these concerns and in addition allow for aggregation across tasks and safety bounds and an open-source evaluation suite to support the reliable characterization of safety in future safe RL research is provided.
Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answers with a filter-then-clone pipeline: a state-only value tr...
An optimal and convergent model-free policy gradient (PG) reinforcement learning (RL) reinforcement learning framework for controlling nonlinear dynamical systems under hard safety constraints is presented and a model-free PG algorithm based on stochastic gradient ascent is developed.
Vipul K. Sharma, Wesley A. Suttle, S. Sivaranjani· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.