Skip to content
Preprint

When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents

Aug 2026 · 3 citations · 27 references
Computer Science

TL;DR

State-Matched Routing and Contextualized Self-Distillation (SMRC-SD), which explicitly determines when and how a privileged trajectory should guide an on-policy student, consistently outperforms unconditional successful full-path distillation.

Abstract

Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only references, such as successful trajectories. In interactive environments, however, the student's preceding actions continually change the execution state. As the student takes different actions or completes subgoals in a different order, its rollout may reach states not covered by the reference, making the reference an unreliable source of guidance for the state actually reached. Applying privileged distillation indiscriminately therefore creates state--reference mismatch. This mismatch motivates a central objective: providing privileged reference guidance that remains compatible with the student's current execution state. We introduce State-Matched Routing and Contextualized Self-Distillation (SMRC-SD), which explicitly determines when and how a privileged trajectory should guide an on-policy student. At each turn, SMRC-SD verifies whether the student's current execution state matches a supported state along the reference trajectory. Distillation is applied only at matched states, filtering out turns for which the reference lacks locally compatible guidance. For each matched state, SMRC-SD further constructs state-conditioned teacher context from the successful trajectory, grounding supervision in the state actually reached. Across ALFWorld and WebShop, SMRC-SD consistently outperforms unconditional successful full-path distillation. With Qwen3-1.7B, it improves task success from $0.746$ to $0.865$ on ALFWorld and from $0.574$ to $0.693$ on WebShop. Controlled routing and context ablations support both selecting locally supported turns and constructing state-compatible teacher context as contributors to these gains. Code is available at https://github.com/liujunzhuo/SMRC-SD.

View source

Similar papers

#machine learning Preprint Sep 2026

Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning

On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factor...

Md. Ismail Hossain, Humaira Kousar, Isidora Chara Tourni · 0 citations
Preprint Aug 2026

PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation

Privileged Adaptation from Student Trajectories (PAST), which treats each completed student trajectory as additional privileged information for the OPSD teacher while leaving the student's distillation prefixes unchanged, and improves the Avg@12 macro average over Vanilla OPSD by 5.6 percentage points.

Yang-Yang Feng, Zhuoyan Feng, Jun-Lan Chen · 5 citations
#artificial intelligence Preprint Sep 2026

TISD: On-Policy Self-Distillation with Trajectory Intervention

A simple branch-regenerate-distill algorithm, Trajectory-Intervention Self-Distillation (TISD), which forces a teacher-selected branch action, returns suffix generation to the student, and distills the full trajectory under the privileged-context-conditioned teacher, support teacher-guided branching as a way to expose...

Taeckyung Lee, Rinat Amankos, Jeonghye Kim et al. · 0 citations
#natural language process... Preprint Sep 2026

USA: Update-aware SAM for Cross-domain On-Policy Disitllation of Language Agents

On-policy distillation instils multi-turn agentic reasoning through dense token-level supervision on the student's own trajectories, but a single domain saturates early, so further supervision has to be drawn from other domains. Multi-domain data mixing is the most direct way of incorporating them, at the cost of confl...

Qi-Yong Zhong, Mao Zheng, Ming-Yang Song et al. · 1 citation
#artificial intelligence Preprint Sep 2026

ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents

Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privil...

Yong Du, Tong-I Chen, Zhengxi Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

IGSD: Environment-Verified Hindsight Self-Distillation for Search Agents

On-policy self-distillation densifies agent training without external teachers: a policy conditioned on privileged hindsight provides step-level guidance for its own unprivileged rollouts. For search agents, however, hindsight can make the teacher prefer a query that does not improve retrieval from the student's state....

Ang-Qing Jiang, Gao-Ming Zhang, Chao-Qun Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.