Preprint
Jul 2026
Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems
PEARL employs an actor-adjoint algorithm that leverages automatic differentiation to compute policy gradients over short horizons and adjoint-based sensitivities of future returns approximated via neural networks, significantly reducing the number of environment interactions, while mitigating long-term gradient instabilities.
Matteo Tomasetto, Nicolò Botteghi, Gabriele Bruni et al.
· 0 citations