Preprint
Jul 2026
Distilled Reinforcement Learning for LLM Post-training
Extensive experiments show that Distilled RL substantially outperforms standard RL and OPD in terms of both pass@1 and pass@k, and can effectively transfer previously unavailable knowledge from a teacher model to a student model.
Chen Wang, Zhaochun Li, Jionghao Bai et al.
· 2 citations