Sep 2026· Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence· 0 citations· 44 references
TL;DR
A probabilistic policy fixing framework that adapts norm-agnostic policies online and provides guarantees that fixed policies are near optimal, given a specified level of confidence is presented.
Abstract
Reinforcement learning (RL) is commonly used to learn reward-optimizing policies. However, RL policies are not always trained with ethical behavior in mind, which can lead an agent to violate social or legal norms in pursuit of its goal. Retraining agents with additional norms is not always feasible, especially in complex stochastic environments. To mitigate this issue, we present a probabilistic policy fixing framework that adapts norm-agnostic policies online. Using Answer Set Programming (ASP), we generate policy fixes that minimize deviations from the RL policy while optimizing for norm adherence against a set of sampled worlds. Based on the Rule of Three and Hoeffding's inequality, we provide guarantees that fixed policies are near optimal, given a specified level of confidence.
Evaluation metrics for safe RL are introduced that address each of these concerns and in addition allow for aggregation across tasks and safety bounds and an open-source evaluation suite to support the reliable characterization of safety in future safe RL research is provided.
This work formalizes the hybrid LLM-planner and RL-controller architecture as a Goal-Augmented Markov Decision Process and shows that when the LLM per-state progress score is used as a bounded potential function, the resulting shaping term preserves the optimal policy set even when the LLM scores are inaccurate.
Christophe D. Hounwanou, John Emeka Eze, Yaé Ulrich Gaba· 0 citations
This paper proposes a safe meta-RL framework that explicitly accounts for safety during adaptation, and develops a safe meta-RL algorithm that learns the safety value function and leverages it for safety filtering and constrained policy optimization.
Exploration in reinforcement learning (RL) remains a fundamental challenge. Recent goal-conditioned RL strategies (which select goals to encourage broader state coverage) have shown promising results, but none scores a goal by novelty and reachability jointly: the two signals are traded off by hand, applied in sequence...
Wen-Yan Yang, A. Mustafin, Dominik Baumann et al.· 0 citations
Reinforcement learning (RL) has become an effective post-training paradigm for long-horizon large language model (LLM) agents. However, we find that the resulting policies can be sensitive to various policy perturbations, such as hidden-state noise, pruning, and quantization. In this work, we study how to improve pertu...
Peng-Xin Wang, Yuan-Zhe Li, Yuxin Ren et al.· 0 citations
A framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator is developed, yielding a persistent safety set from which the agent can remain safe indefinitely, and a new reward maximization algorithm is proposed that effectively exploits the learned persistent safety set for rewar...
A. Choudhury, J. Brahmanage, Akshat Kumar et al.· Proceedings of the Thirty-Fi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.