When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning
This work introduces OVI, an interactive on-policy IL algorithm that is statistically efficient whenever the learner can represent the expert's value function and computationally efficient given access to a linear maximization oracle, and introduces a negative result showing that interaction is necessary.