Skip to content

Evaluation Metrics for Safe Reinforcement Learning

Sep 2026 · 0 citations · 24 references
Computer Science

TL;DR

Evaluation metrics for safe RL are introduced that address each of these concerns and in addition allow for aggregation across tasks and safety bounds and an open-source evaluation suite to support the reliable characterization of safety in future safe RL research is provided.

Abstract

Safe reinforcement learning (RL) is commonly formalized as a Constrained Markov Decision Process (CMDP), in which an agent maximizes expected reward while keeping its expected cumulative cost below a specified safety bound. Existing safe RL benchmarks predominantly report whether an algorithm is safe on average, following this expectation-based guarantee. We argue that this convention is insufficient to reliably characterize an algorithm's true safety: it fails to capture how often and how severely the safety bound is violated, whether this holds consistently across tasks and safety bounds, and whether training-time behavior is representative of behavior of the final converged policy. Therefore, we introduce (i) evaluation metrics for safe RL that address each of these concerns and in addition allow for aggregation across tasks and safety bounds. We furthermore define (ii) a safety tier system to systematically categorize and compare algorithms in terms of safety and reliability at both training and for a final policy. Using this framework, we provide (iii) an empirical safety evaluation across multiple safety navigation tasks. Our results show that aggregate metrics, distributional reporting, and task- and safety bound-specific results each reveal information the other metrics cannot. We therefore recommend reporting all three jointly, rather than compressing this information into a single value, as is common practice. We provide SafeRLEval, an open-source evaluation suite to support the reliable characterization of safety in future safe RL research.

View source

Similar papers

Conference Open access Sep 2026

Persistent Safety Set Guided Offline Safe Reinforcement Learning

A framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator is developed, yielding a persistent safety set from which the agent can remain safe indefinitely, and a new reward maximization algorithm is proposed that effectively exploits the learned persistent safety set for rewar...

A. Choudhury, J. Brahmanage, Akshat Kumar et al. · 0 citations
#machine learning Preprint Sep 2026

Safe Meta-Reinforcement Learning via Information Space Reachability

This paper proposes a safe meta-RL framework that explicitly accounts for safety during adaptation, and develops a safe meta-RL algorithm that learns the safety value function and leverages it for safety filtering and constrained policy optimization.

Ze-Yang Li, Sunbochen Tang, Navid Azizan · 0 citations
Conference Open access Sep 2026

ASP-Based Probabilistic Policy Fixing for Norm Compliant RL

A probabilistic policy fixing framework that adapts norm-agnostic policies online and provides guarantees that fixed policies are near optimal, given a specified level of confidence is presented.

Sebastian P. Adam, Thomas Eiter · 0 citations
#machine learning Preprint Sep 2026

Certified Safety Curation: Distribution-Free Guarantees for Safe Offline Reinforcement Learning

Safe offline reinforcement learning assumes a cost function on every transition. We ask what remains possible when safety can be judged only by comparing short clips and occasionally asking whether an episode exceeded its budget. Certified safety curation answers with a filter-then-clone pipeline: a state-only value tr...

Adam Haroon, Cody H. Fleming · 0 citations
#reinforcement learning Open access Aug 2026

Learning from the Test: Self-Referential Differential Testing for Deep RL Agents

Delta (Differential Testing for DRL Agents) is proposed, a novel and comprehensive framework that automatically identifies both safety-critical and optimality bugs in DRL agents and investigates the effectiveness of three offline RL algorithms in generating challenger agents.

Jun-Da He, Jie-Ke Shi, Zhou Yang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SUN: Reaching for Novelty in Reinforcement Learning

Exploration in reinforcement learning (RL) remains a fundamental challenge. Recent goal-conditioned RL strategies (which select goals to encourage broader state coverage) have shown promising results, but none scores a goal by novelty and reachability jointly: the two signals are traded off by hand, applied in sequence...

Wen-Yan Yang, A. Mustafin, Dominik Baumann et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.