Aug 2026· IEEE/ASME International Conference on Mechatronic and Embedded Systems and Applications· pp. 33-38· 0 citations· 24 references
Abstract
Offline safe reinforcement learning aims to learn constraint-satisfying policies from pre-collected datasets without online interaction, which is critical for safety-critical applications such as autonomous driving. However, offline datasets collected from heterogeneous sources often contain spurious correlations between cost-irrelevant state features and safety signals, which can mislead policy learning and cause constraint violations under distribution shift. To address this issue, we propose a two-stage framework: (1) a causal model is employed to identify state features that are necessary and sufficient for safety constraints; (2) a safe policy is trained on the extracted features using Hamilton-Jacobi reachability-based cost value estimation and gradient-based action correction. Experiments on the MetaDrive autonomous driving benchmark demonstrate that our method achieves state-of-the-art performance in both reward and cost constraint satisfaction, and ablation study confirms that causal feature extraction significantly improves policy performance and training stability in complex driving scenarios. This work highlights the promise of causal representation learning as a principled approach to improving both performance and safety in offline reinforcement learning.
A framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator is developed, yielding a persistent safety set from which the agent can remain safe indefinitely, and a new reward maximization algorithm is proposed that effectively exploits the learned persistent safety set for rewar...
A. Choudhury, J. Brahmanage, Akshat Kumar et al.· Proceedings of the Thirty-Fi...· 0 citations
This paper addresses the problem of safe offline reinforcement learning, which involves training a policy to satisfy safety constraints using an offline dataset. This problem is inherently challenging as it requires balancing three highly interconnected and competing objectives: satisfying safety constraints, maximizin...
Sheng-Chao Hu, Peng Wang, Ji-Feng Hu et al.· 0 citations
Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains challenging due to insufficient exposure to safety-critical driving scenes. The long-tailed nature of real-world traffic situations makes dangerous and rare interactions difficult to...
Xincong Hu, Lei Ou, Mao-Sen Li et al.· 0 citations
Robots deployed for competitive tasks must outmaneuver their opponents without sacrificing safety. Existing approaches, including safe reinforcement learning (RL), train a single policy to achieve task success and avoid failures simultaneously. This coupling can complicate training and leave the learned policy exploita...
Rui-Han Wu, Rui Yang, Donggeon David Oh et al.· 0 citations
Autonomous driving has the potential to greatly enhance traffic efficiency, and its effectiveness depends on robust decision-making in complex real-world environments. As an emerging technique, Deep Reinforcement Learning (DRL) is expected to address this requirement. However, most existing general-purpose DRL methods...
Rui Guo, Xin-Yu Li, Zhong-Hao Fu et al.· IEEE Open Journal of Intelli...· 0 citations
SAGE improves robustness in novel and safety-critical scenarios, reduces safety violations, and maintains task performance comparable to strong baseline policies, suggesting that agents can improve after initial training by estimating the limits of their competence, requesting guidance when needed, and learning selecti...
Dong Hu, Chao Huang, Carman K. M. Lee et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.