Skip to content

Flux-OPD: On-Policy Distillation with Evolving Contexts

Jul 2026 · arXiv.org · Vol abs/2607.28022 · 1 citation · 27 references
Computer Science

TL;DR

Flux-OPD is proposed, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains and outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.

Abstract

Large language model training in open-ended domains lacks verifiable rewards, making task preferences difficult to formalize as effective supervision. Contexts can convey such preferences, yet provide little additional supervision once distilled into the student, motivating contexts that evolve with student performance. However, directly using evolving contexts as in-training supervision results in an unstable distillation target and conflicting distributions, requiring mechanisms to stabilize target and downweight conflicts. In this paper, we analyze the effect of contexts through a decomposition of the reverse KL objective, revealing two findings: the student is distilled toward the geometric mean of context-conditioned teachers, and the objective contains a conflict term that measures conflicts among these teachers. Based on this decomposition, we propose Flux-OPD, an OPD paradigm that uses evolving contexts as in-training supervision to capture task preferences in open-ended domains. Flux-OPD treats the differences between context-conditioned and context-free teachers as contextual difference signals, injects them as contextual corrections into the context-free teacher anchor, and weights their correction strength using the conflict term as an indicator. Experiments on open-ended tasks show that Flux-OPD outperforms existing OPD paradigms, highlighting the potential to combine teacher supervision with evolving contexts.

View source

Similar papers

Preprint Aug 2026

Adaptive Supervised Anchoring for On-Policy Self-Distillation

Context quality is identified as a central bottleneck in on-policy self-distillation and the value of separating rollout-conditioned guidance from canonical supervision is demonstrated, demonstrating the value of separating rollout-conditioned guidance from canonical supervision.

Meilin Yang, Zixuan Ding, Jianhao Nie et al. · 0 citations
Preprint Aug 2026

Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation

The role of target-specific privilege is investigated with On-Policy Self-Distillation from Other Problems (On-Policy Self-Distillation from Other Problems), which replaces the paired reference with a problem and solution from a different example, while preserving the student rollout, teacher, and distillation objectiv...

Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar et al. · 9 citations · ⚡3
Preprint Aug 2026

DAPD: Dual-Anchored Policy Distillation

DAPD is proposed, a unified framework with two levels of anchoring that significantly alleviates privilege illusion, outperforming OPSD on Qwen3-4B by +2.00 points on average across tasks.

Jian-Yu Wu, Yi-Zhou Wang, Encheng Su et al. · 1 citation
Review Aug 2026

One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation

This review treats collapse as a symptom governed by three levers: where the signal is applied, that is, how tokens are weighted; what the teacher is shown, that is, the nature of the privileged information; and when the signal changes, that is, the teacher's dynamics and the decay of guidance.

J. Robert, Raheel Qader · 0 citations
#artificial intelligence Preprint Sep 2026

Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation

On-policy distillation (OPD) is a promising approach for transferring knowledge between language models, where a student receives dense token-level supervision along its own generated trajectories. However, teacher supervision can be unreliable when conditioned on incomplete or low-quality student prefixes. We identify...

Jin-Gang Zhou, Yu-Yi Zhou, Hai-Yang Guo et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.