Skip to content

Self-Play Search Distillation for Large Language Model Reasoning

Sep 2026 · 0 citations · 59 references
Computer Science

TL;DR

Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games, offers an annotation-efficient way to create high-quality synthetic data for improving LLM performance in reasoning tasks.

Abstract

Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games. SPSD uses executable environments to turn search into structured reasoning problems. At each state, the expert identifies a preferred decision, plausible alternatives, plausible opponent replies, and value estimates. By converting the self-play search records into superhuman chains-of-thought, we train LLMs with environment-grounded supervision. Although trained only on self-play search records, SPSD transfers to unseen mathematics. On Qwen3-4B-Base, it raises the mean over six mathematics benchmarks from 24.1 to 36.6 while increasing the held-out-game win rate from 15% to 45%. SPSD offers an annotation-efficient way to create high-quality synthetic data for improving LLM performance in reasoning tasks.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

OPSRD: On-Policy Self-Role Distillation

Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transf...

Wei-Jie Ren, Yan-Wen Zhang, Hao Li et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models

SOLID is proposed, a novel framework for self-improving OR language models without verified answers or external evaluators that improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training.

Rui-Chen Zhu, Ming-Long Cao, Chen-Yu Zhou et al. · 0 citations
Preprint Aug 2026

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

This work shows that reasoning traces are not a prerequisite: skills distilled from non-reasoning trajectories alone remain competitive with skills distilled from paired reasoning/non-reasoning corpora, with domain-dependent differences between the two sources.

Agamdeep Singh, Srishti Gautam, Priyanshu Gupta et al. · 0 citations
#machine learning Preprint Sep 2026

Negative Self-Distillation: Learning to Reason by Avoiding Flaws

On-Policy Self-Distillation (OPSD) has emerged as a popular paradigm for large language model (LLM) self-improvement, allowing models to act as their own teachers by leveraging privileged information such as ground-truth solutions. However, recent findings indicate that OPSD can severely degrade the performance of LLMs...

Rong-Can Pei, Zhepei Wei, Shu-Yao Xu et al. · 1 citation
#artificial intelligence Preprint Sep 2026

Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces

Large language models (LLMs) excel at reasoning when scaled to hundreds of billions of parameters, but small- and mid-scale models remain brittle reasoners even with knowledge distillation (KD). We present Ladders-of-Thought (LoT), a framework that improves reasoning by combining progressive question rewrites with a se...

Ming-Hui Liu, Thomas Magelinski, De-Hao Yuan et al. · 0 citations
#natural language process... Preprint Sep 2026

Learning from Think-Mode Advantage via On-Policy Distillation

Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes. Privileged reaso...

Wan-Qi Ren, Jian-Xiang Wang, Dan-Xuan Liu et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.