Skip to content

Robust Adversarial Reinforcement Learning with Risk Sensitivity and Critic Consistency Regularization

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

This work proposes a unified framework, RACER (Risk-sensitive robust Adversarial critic ConsistEncy-regularized Reinforcement learning), that revisits adversarial RL from a risk-sensitive perspective and introduces a state-dependent adversarial objective that adaptively regulates perturbation strength.

Abstract

Reinforcement learning (RL) achieves strong performance in sequential decision-making but remains brittle under dynamic uncertainty and distributional shifts. Robust Adversarial Reinforcement Learning (RARL) improves robustness via worst-case perturbations, but existing approaches frequently suffer from unstable optimization and degraded value estimation. In particular, overly aggressive adversaries can drive the agent toward uninformative failure states, while adversarial perturbations amplify disagreement between double critics and introduce biased value targets. We propose a unified framework, RACER (Risk-sensitive robust Adversarial critic ConsistEncy-regularized Reinforcement learning), that revisits adversarial RL from a risk-sensitive perspective. First, we introduce a state-dependent adversarial objective that adaptively regulates perturbation strength, suppressing harmful disturbances while preserving informative exploration. Second, we propose critic consistency regularization to reduce disagreement between Q-value estimators and stabilize learning. Comprehensive experiments on challenging continuous control benchmarks demonstrate that RACER consistently improves performance, robustness, and training stability over strong robust RL baselines.

View source

Similar papers

#machine learning Preprint Aug 2026

RL-FAT: Reinforcement Learning for Fair Adversarial Training

RL-FAT is proposed, a reinforcement-learning-inspired fair adversarial training framework that uses policy-gradient based feedback from adversarial predictions to improve adversarial robustness while promoting a more balanced robustness distribution across classes.

Tejaswini Medi, Levan Mikeladze, Margret Keuper · 0 citations
Conference Aug 2026

Robust Deep Reinforcement Learning via Adversarial Data Augmentation for Robotic Control Under Observation Uncertainty

Observation uncertainty can drive a deep reinforcement learning (DRL) controller to select inappropriate actions even when its nominal policy performs well. This issue is especially pronounced in robotic systems deployed outside controlled laboratory conditions, where sensing and communication are affected by noise, bi...

Hai-Tao Zhang, Cheng-Yu Liu, Qing-Kai Meng et al. · 0 citations
#machine learning Preprint Sep 2026

Robust Policy Optimization via Adversarial Importance Sampling

This work introduces Adversarial Importance Sampling (Advis), a method that uses importance sampling over trajectories from standard training to estimate and optimize verifiable worst-case returns and introduces advrl, a modular PyTorch library that provides clean, single-file implementations of existing robustness met...

Amine Andam, Jamal Bentahar, M. Hedabou · 0 citations
#machine learning Preprint Sep 2026

Certifying Lower Bounds for Risk-Sensitive Reinforcement Learning under Adversarial State Perturbations

This paper proposes an empirical method that improves certified lower bounds by selecting the training risk-aversion parameter $\beta$ independently of the risk level used during evaluation, and formulate the risk-sensitive certification problem as a convex optimization and derive its dual to obtain a tractable approxi...

Tong Li, Saunak Kumar Panda, Yi-Sha Xiang · 0 citations
#machine learning Preprint Sep 2026

Admissable: Training Reinforcement Learning Agents against Adversarial Missingness

In order to make Reinforcement Learning algorithms applicable in real world scenarios, safety must be ensured even under adverse operating conditions. In this work, we consider the challenge of adversarial feature missingness: a scenario in which an adversary occludes features from the agent's observation in order to r...

Paul Stahlhofen, Luca Hermes, Tim Kochs et al. · 0 citations
#machine learning Preprint Sep 2026

Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning

Dually Regularized AIL is the first algorithm to simultaneously achieve $\widetilde{O}\left(\frac{1}{\epsilon}\right)$ sample complexity in both expert demonstrations and online interactions for this regularized AIL objective, even with stochastic experts.

Hang Zhou, Shang-Zhe Li, А. А. Браверман et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.