Skip to content

SCOPE-OPSD: Fisher-Conditioned Privileged Subspaces for On-Policy Self-Distillation

Sep 2026 · 1 citation · 34 references
Computer Science

TL;DR

The results support a compact, Fisher-conditioned privileged subspace for short-budget OPSD, with strict gains in 11 of the 12 model-checkpoint combinations and an exact tie at 4B step 25.

Abstract

On-policy self-distillation (OPSD) scores student-generated prefixes with a solution-conditioned self-teacher, yet transfers supervision only through next-token probabilities. We ask whether the aligned final-layer discrepancy offers a useful second channel, and how to test that channel without confusing its geometry with auxiliary strength. SCOPE-OPSD projects the privileged teacher-student residual onto a frozen rank-64 factor estimated from residual covariance and language-model-head Fisher sensitivity. It reuses the forwards already required by OPSD and adds neither rollouts nor inference-time modules. A matched Random control preserves the structured factor's rank and nonzero spectrum and uses per-arm gradient-RMS calibration, isolating the effect of the data-dependent orientation. Across the complete 25/50/75/100-step trajectories for Qwen3-1.7B, 4B, and 8B, Structured is never below Pure OPSD, with strict gains in 11 of the 12 model-checkpoint combinations and an exact tie at 4B step 25. Structured also exceeds matched Random in 10 of the 12 combinations. At step 75 on Qwen3-1.7B, Structured exceeds matched Random by 1.39 Macro Avg@12 points in each of two independent training reruns. A cross-fitted diagnostic also shows 4.40 times greater held-out privileged-gap capture than the matched random orientation. The results support a compact, Fisher-conditioned privileged subspace for short-budget OPSD.

View source

Similar papers

Preprint Aug 2026

Adaptive Supervised Anchoring for On-Policy Self-Distillation

Context quality is identified as a central bottleneck in on-policy self-distillation and the value of separating rollout-conditioned guidance from canonical supervision is demonstrated, demonstrating the value of separating rollout-conditioned guidance from canonical supervision.

Meilin Yang, Zixuan Ding, Jianhao Nie et al. · 0 citations
Preprint Aug 2026

Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

Self-OPD is introduced, a teacher-free OPD framework for flow matching models that turns the student's own self-exploration into step-wise supervision and outperforms prior RL and OPD methods without task-specific teachers.

Shi-Yi Zhang, Mu-Shui Liu, Yunze Tong et al. · 1 citation
#machine learning Preprint Aug 2026

SR-OPSD: Self-Referenced On-Policy Self-Distillation

Extensive experiments across scientific evaluation, mathematical reasoning, and coding generation tasks with multiple large language models show that SR-OPSD achieves the state-of-the-art or competitive performance across various settings.

Zhuo Sun, Entong Li, Yan-Long Zhao et al. · 0 citations
#machine learning Preprint Sep 2026

Activation-Conditioned Self-Distillation

On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distilla...

Zhe-Xi Lu, Subhajit Chaudhury, Tejaswini Pedapati et al. · 0 citations
Preprint Aug 2026

WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training

WDL-OPD is introduced, a mixture-constrained co-training method with two trainable policies that shows that freezing the auxiliary recovers an anchor-plus-contrast proxy target closely related to OPD$^2$ and W2S-OPD, whereas joint training creates branch-level degrees of freedom that a static delta cannot express.

Ze-Hao Chen, Gong-Xun Li, Tianxiang Ai et al. · 0 citations
#machine learning Preprint Sep 2026

Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factor...

Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.