Self-Play Search Distillation for Large Language Model Reasoning
Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games, offers an annotation-efficient way to create high-quality synthetic data for improving LLM performance in reasoning tasks.