Skip to content

Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models

Sep 2026 · 0 citations · 33 references
Mathematics Computer Science

TL;DR

SOLID is proposed, a novel framework for self-improving OR language models without verified answers or external evaluators that improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training.

Abstract

Modern large language models (LLMs) can translate natural-language descriptions into operations research (OR) formulations. Post-training techniques including reinforcement learning and on-policy self-distillation have further improved this capability. However, three limitations remain in training LLMs for OR formulations. First, training commonly relies on synthetic formulations validated by human experts or stronger models, constraining scalable supervision. Second, credit assignment is either coarse or costly: outcome rewards score an entire trajectory without locating the responsible modeling decision, whereas process-level supervision requires an additional evaluator. Third, privileged self-distillation can induce style mismatch by using solver context unavailable at deployment. We find that a model can improve from solver-artifact feedback generated by its own rollouts, making self-distillation a practical, evaluator-free source of dense supervision. Therefore, we propose SOLID: Solver-Informed On-Policy LearnIng through Self-Distillation, a novel framework for self-improving OR language models without verified answers or external evaluators. SOLID executes candidate programs from multiple rollouts, clusters their objectives, and selects a majority-group artifact as a pseudo-reference. The model then performs updates using group-relative advantages and dense self-supervision signals. Across multiple OR benchmarks, SOLID improves solution accuracy for both general-purpose and OR-tuned models over outcome-only group-relative training. These results show that solver artifacts can support scalable self-improvement without trusted answers.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs

Multi-Teacher Self-Distillation Policy Optimization is introduced, an on-policy distillation method that unifies several frozen teachers into one student model and lifts the weakest domain of Qwen3-8B by 14.79 points and narrows its domain gap by 74.7%, a better balance than serving one matched teacher per domain.

Xi-Xiang He, Xing-Ming Li, Bai-Qi Wu et al. · 2 citations
#machine learning Preprint Sep 2026

Learning to Optimize through Solver-Grounded Self-Play

Results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.

Xia Jiang, Yao-Xin Wu, Chen-Yu Zhou et al. · 0 citations
Preprint Aug 2026

On-Policy Self-Distillation without Any Supervision

U-OPSD first samples multiple rollouts and constructs a pseudo solution by majority vote under a self-consistency threshold, and conditions the model's distribution on the pseudo-solution and distills itself on the disagreeing completions, allowing the model to correct itself precisely where it is confidently wrong.

Yijiang Li, Bingyang Wang, Yijun Liang et al. · 10 citations · ⚡1
#artificial intelligence Preprint Sep 2026

Self-Play Search Distillation for Large Language Model Reasoning

Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games, offers an annotation-efficient way to create high-quality synthetic data for improving LLM performance in reasoning tasks.

Lorenzo Molfetta, Wai-Chung Kwan, Giacomo Frisoni et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.