Preprint
Jul 2026
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning
This work introduces OVI, an interactive on-policy IL algorithm that is statistically efficient whenever the learner can represent the expert's value function and computationally efficient given access to a linear maximization oracle, and introduces a negative result showing that interaction is necessary.
Luca Viano, Antoine Moulin, Audrey Huang et al.
· 0 citations