Accelerating Visual On-Policy Distillation with Batched Speculative Jacobi Rollouts
Visual on-policy distillation (OPD) improves the training of compact visual autoregressive models by learning from trajectories generated by the current student. However, these online rollouts are still produced token by token with autoregressive decoding, which adds substantial cost to every on-policy training step. S...