Skip to content

Counterfactual Constraint-Conditioned On-Policy Distillation for Multi-Constraint Instruction Following

Sep 2026 · 0 citations · 25 references
Computer Science Mathematics

TL;DR

CC-OPD (Counterfactual Constraint-Conditioned On-Policy Distillation), which inverts the standard supervision-generation direction in distillation, and achieves the highest average among all evaluated student-training methods.

Abstract

Multi-constraint instruction following requires a model to respond to a query under many simultaneously active constraints. Even strong instruction-tuned models still routinely violate some of them. Existing approaches either augment supervision with sequence- or token-level RL rewards from external verifiers or learned graders, or use on-policy distillation (OPD) against a single full-context teacher whose probability mass becomes diluted as more constraints become simultaneously active. We propose CC-OPD (Counterfactual Constraint-Conditioned On-Policy Distillation), which inverts the standard supervision-generation direction in distillation. Rather than enriching the teacher with information beyond what the student sees, CC-OPD ablates each constraint from the teacher's conditioning in turn, and constructs the per-constraint signal from the resulting per-token probability differentials. The resulting per-token leave-one-out log-likelihood shifts are summed, clipped, and added to the vanilla OPD reward as a token-level shaping term. All shaping terms are obtained from the frozen teacher, without an external verifier during distillation, and the reward equals vanilla OPD wherever the aggregate shift is zero. Across two Qwen model pairs and seven benchmarks, CC-OPD achieves the highest average among all evaluated student-training methods. A 1.5B student trained with CC-OPD surpasses its own 7B RL-trained teacher on the MulDimIF benchmark.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

Group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance, and demonstrates that group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance.

Zhu Zhang, Ji-Xun Wang, Xiao-An Xu et al. · 2 citations
Preprint Aug 2026

WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training

WDL-OPD is introduced, a mixture-constrained co-training method with two trainable policies that shows that freezing the auxiliary recovers an anchor-plus-contrast proxy target closely related to OPD$^2$ and W2S-OPD, whereas joint training creates branch-level degrees of freedom that a static delta cannot express.

Ze-Hao Chen, Gong-Xun Li, Tianxiang Ai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beam Search as Test-Time Self-Distillation via Counterfactual Contexts

Self-Distillation Fine-Tuning (SDFT) enables a language model to act as its own teacher: by conditioning on a demonstration, the model produces an implicit reward via pointwise mutual information, which guides on-policy learning without external supervision. However, SDFT operates at training time: it requires gradient...

Su-Ee Tan, Xiao-Tong Ji, Rasul Tutunov et al. · 0 citations
#artificial intelligence Preprint Sep 2026

OPSRD: On-Policy Self-Role Distillation

Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transf...

Wei-Jie Ren, Yan-Wen Zhang, Hao Li et al. · 0 citations
#natural language process... Preprint Sep 2026

Learning from Think-Mode Advantage via On-Policy Distillation

Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-policy distillation (OPD). OPD preserves student-generated trajectories and provides dense token-level teacher targets at student-visited prefixes. Privileged reaso...

Wan-Qi Ren, Jian-Xiang Wang, Dan-Xuan Liu et al. · 0 citations

Related blog posts

GPT-Lab Sep 3, 2026

Adaptive AI Agents in Construction Workflows

Adaptive AI agents can help make BIM data more machine-readable by navigating IFC models, interpreting inconsistent information, and mapping it to defined standards. In this blog, Alok Rawat shares findings from a real-world pilot in construction workflows. The post Adaptive AI Agents in Construction Workflows appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.