Off-Policy Inverse Q-Learning Algorithm with Unobservable States
To address the problem of an unknown performance index in discrete-time (DT) linear systems with unobservable states, this paper investigates the output-feedback inverse reinforcement learning (IRL) problem and proposes a data-driven output-feedback off-policy inverse Q-learning algorithm. The proposed method relies so...