MInTRL: Off-policy Intervention can boost On-policy RL
This work introduces Minimal Intervention Reinforcement Learning (MInTRL), which expands the exploration frontier through sparse, local interventions in otherwise on-policy rollouts, and establishes minimal intervention as an effective paradigm for enhancing on-policy RL.
Ming-Yu Chen, Ye-Fan Tao, Gerald Friedland et al.
· 0 citations