Skip to content

Robust Policy Optimization via Adversarial Importance Sampling

Sep 2026 · 0 citations · 36 references
Computer Science

TL;DR

This work introduces Adversarial Importance Sampling (Advis), a method that uses importance sampling over trajectories from standard training to estimate and optimize verifiable worst-case returns and introduces advrl, a modular PyTorch library that provides clean, single-file implementations of existing robustness methods and adversarial attacks.

Abstract

Significant progress has been made in safeguarding deep reinforcement learning (DRL) policies against input perturbations. Developing robust DRL involves three main stages: algorithm design, implementation, and evaluation. In this work, we identify and address a key limitation at each stage. First, we introduce Adversarial Importance Sampling (Advis), a method that uses importance sampling over trajectories from standard training to estimate and optimize verifiable worst-case returns. Advis satisfies three desirable criteria not jointly achieved by prior work: it requires no additional environment interactions, no auxiliary networks, and captures long-term robustness. Second, we introduce advrl, a modular PyTorch library that provides clean, single-file implementations of existing robustness methods and adversarial attacks, facilitating rapid prototyping and enabling reproducible and traceable evaluations. Third, we revisit evaluation under learned adversaries and show that optimal adversarial hyperparameters do not transfer across agents, which can lead to an overestimation of robustness when using a limited set of attacker configurations. Accordingly, we evaluate policies against a large and diverse set of attackers, using 6-14x more configurations than prior work. Finally, we evaluate our approach on continuous control environments, demonstrating its effectiveness relative to existing baselines. The code is available at: https://github.com/AmineAndam04/advrl

View source

Similar papers

Conference Aug 2026

Robust Deep Reinforcement Learning via Adversarial Data Augmentation for Robotic Control Under Observation Uncertainty

Observation uncertainty can drive a deep reinforcement learning (DRL) controller to select inappropriate actions even when its nominal policy performs well. This issue is especially pronounced in robotic systems deployed outside controlled laboratory conditions, where sensing and communication are affected by noise, bi...

Hai-Tao Zhang, Cheng-Yu Liu, Qing-Kai Meng et al. · 0 citations
#machine learning Preprint Sep 2026

Admissable: Training Reinforcement Learning Agents against Adversarial Missingness

In order to make Reinforcement Learning algorithms applicable in real world scenarios, safety must be ensured even under adverse operating conditions. In this work, we consider the challenge of adversarial feature missingness: a scenario in which an adversary occludes features from the agent's observation in order to r...

Paul Stahlhofen, Luca Hermes, Tim Kochs et al. · 0 citations
#machine learning Preprint Sep 2026

Robust Adversarial Reinforcement Learning with Risk Sensitivity and Critic Consistency Regularization

This work proposes a unified framework, RACER (Risk-sensitive robust Adversarial critic ConsistEncy-regularized Reinforcement learning), that revisits adversarial RL from a risk-sensitive perspective and introduces a state-dependent adversarial objective that adaptively regulates perturbation strength.

Jia-Xin Wu, Tian-Tian Zhang, Yu-Xing Wang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Neither Adversarial Training Nor Purification: Emergent Adversarial Robustness from Oscillatory Predictive Learning

This work introduces Oscillatory Predictive Learning (OPL), a two-stage framework that combines Artificial Kuramoto Oscillatory Neurons (AKOrN) with predictive self-supervised pretraining using X-PhiNet and compares it with other randomized adversarial defense methods that provide precise, reproducible, and strong atta...

M. Habibi, Klea Ziu, Martin Takác et al. · 0 citations
Preprint Aug 2026

Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks

This work proposes BMAT (Bilevel-Minimax Adversarial Transfer), an integrated bottom-up solver that combines a Soft Weight Modulator and an Implicit Gradient Approximator to enable ternary coupling among initialization, surrogate adaptation, and perturbation optimization.

Yao-Hua Liu, Yi-Fan Guo, Jia-Xin Gao · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.