Robust Adversarial Reinforcement Learning with Risk Sensitivity and Critic Consistency Regularization
This work proposes a unified framework, RACER (Risk-sensitive robust Adversarial critic ConsistEncy-regularized Reinforcement learning), that revisits adversarial RL from a risk-sensitive perspective and introduces a state-dependent adversarial objective that adaptively regulates perturbation strength.