Last-Iterate Guarantees for Online Reinforcement Learning in Structured Constrained MDPs
This work develops a general, statistically efficient framework for last-iterate convergence in structured Constrained MDPs (CMDPs), and validate the stabilising effect predicted by the theory on a synthetic linear CMDP: the regularised method exhibits stable last-iterate behaviour, whereas its unregularised counterpar...