Preprint
Aug 2026
Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning
A training method for HIL online reinforcement learning for real robots that automatically switches between learning from interventions and on-policy self-improvement, reducing the policy--target-sample gap that otherwise induces execution-time distribution shift.
Zihang Wang, Yishan Wang
· 0 citations