Skip to content

Safe Meta-Reinforcement Learning via Information Space Reachability

Sep 2026 · 0 citations · 18 references
Computer Science Engineering

TL;DR

This paper proposes a safe meta-RL framework that explicitly accounts for safety during adaptation, and develops a safe meta-RL algorithm that learns the safety value function and leverages it for safety filtering and constrained policy optimization.

Abstract

Meta-reinforcement learning (meta-RL) enables agents to adapt to unseen tasks with limited experience. Despite its promise, the application of meta-RL in real-world tasks is hindered by safety requirements, which have been underexplored in prior work. In this paper, we propose a safe meta-RL framework that explicitly accounts for safety during adaptation. Our key insight is to reason about safety in the information space, which captures both the physical state and the agent's belief over the underlying task. Within this space, we introduce a safety value function that measures the probability of the agent avoiding unsafe regions indefinitely. We show that this function satisfies a self-consistency condition and a Bellman equation, which make it learnable via meta-RL. Based on this formulation, we develop a safe meta-RL algorithm that learns the safety value function and leverages it for safety filtering and constrained policy optimization. Experiments on meta-RL benchmarks demonstrate the effectiveness of the proposed method.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

SUN: Reaching for Novelty in Reinforcement Learning

Exploration in reinforcement learning (RL) remains a fundamental challenge. Recent goal-conditioned RL strategies (which select goals to encourage broader state coverage) have shown promising results, but none scores a goal by novelty and reachability jointly: the two signals are traded off by hand, applied in sequence...

Wen-Yan Yang, A. Mustafin, Dominik Baumann et al. · 0 citations
Conference Open access Sep 2026

Persistent Safety Set Guided Offline Safe Reinforcement Learning

A framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator is developed, yielding a persistent safety set from which the agent can remain safe indefinitely, and a new reward maximization algorithm is proposed that effectively exploits the learned persistent safety set for rewar...

A. Choudhury, J. Brahmanage, Akshat Kumar et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Evaluation Metrics for Safe Reinforcement Learning

Evaluation metrics for safe RL are introduced that address each of these concerns and in addition allow for aggregation across tasks and safety bounds and an open-source evaluation suite to support the reliable characterization of safety in future safe RL research is provided.

Lindsay Spoor, A. Plaat, T. Moerland · 0 citations

Declarative Specifications for Efficient and Safe Reinforcement Learning

This dissertation presents a work in safe RL, where agents must also respect safety constraints using pure-past linear-time temporal logic (PPLTL), and presents how to enforce safety constraints using pure-past linear-time temporal logic (PPLTL).

Giovanni Varricchione · 0 citations
#machine learning Preprint Sep 2026

ICMAPE: In-Context Multiagent Pure Exploration

In some multi-agent systems, the quantity to be optimized is not an externally specified reward but the information acquired about unknown properties of the environment as done in active sequential hypothesis testing (ASHT) problems. However, the ASHT literature tends to focus on finite single-agent problems with well-...

Xin-Yi Hu, Alessio Russo, Aldo Pacchiano · 0 citations
Preprint Aug 2026

Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency

While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses a fundamental obstacle to effective RL improvement. In particular, their severe inference latency---which can lead to pauses or jerky movem...

Brian Zhu, Momen Khalil, E. Harrison et al. · 1 citation

Related blog posts

Microsoft Research Blog Sep 30, 2026

Forecasting space weather risks on power grids

Extreme space-weather events can damage power systems on Earth and degrade GPS accuracy and satellite operations. A new machine learning system can predict where damage is likely to occur 30-60 minutes before a storm arrives. The post Forecasting space weather risks on power grids appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.