Preprint
Aug 2026
Efficient Hypergradient Descent for Inverse Reinforcement Learning
This work shows that, at the inner optimum, the Hessian of the inner objective is proportional to the Fisher information matrix of the policy, yielding a structured Fisher-based hypergradient closely related to Natural Hypergradient Descent.
Nikita Sevriukov, A. Barabanova, Uliana Gagarina et al.
· 0 citations