Skip to content
Preprint

Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL

Aug 2026 · 2 citations · 26 references
Computer Science

TL;DR

Influence Calibration for Self-Distillation (ICSD) improves all matched aggregate metrics over trust-only allocation under Group Relative Policy Optimization (GRPO) and Group-in-Group Policy Optimization (GiGPO) and raises cosine compatibility with the RL gradient.

Abstract

On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories. Existing methods allocate this supervision mainly by teacher trust, but trust does not reveal whether emphasizing a token supports the current policy objective. We call this the trust-utility mismatch and introduce Influence Calibration for Self-Distillation (ICSD). For each supervised token, ICSD measures the first-order response of its importance-weighted RL surrogate contribution to a teacher-directed output perturbation. Batch-adaptive calibration converts this non-stationary signal into a bounded allocation weight while preserving the original auxiliary-loss mass within each action turn. These detached weights affect only the distillation loss and require no additional model pass. Across ALFWorld, WebShop, and Search-QA, ICSD improves all matched aggregate metrics over trust-only allocation under Group Relative Policy Optimization (GRPO) and Group-in-Group Policy Optimization (GiGPO), across two model families spanning 1.5B to 7B. At 7B, it reaches 96.1% ALFWorld success and a WebShop score of 93.1. Frozen-batch analyses show that ICSD reduces teacher-supported mass assigned to objective-opposed tokens from 60.1% to 37.8% and raises cosine compatibility with the RL gradient by 0.192. A companion repository is avail- able at https://github.com/lanqz7766/Influence-Calibration-for-On-Policy-Self-Distillation-in-Agentic-RL.

View source

Similar papers

#machine learning Preprint Sep 2026

Distill What You Trust: Reliability-Aware Multi-Teacher On-Policy Distillation

Multi-teacher on-policy distillation allows a student to learn from complementary specialists on its own trajectories. Domain-routed approaches, however, select one teacher per example and keep it fixed throughout the response. This design both depends on domain labels that mixed training corpora often lack and cannot...

Jie Sun, Mao Zheng, Ming-Yang Song et al. · 3 citations
#artificial intelligence Preprint Sep 2026

When and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment

UECR-GRPO is introduced, which integrates verifier and teacher signals within a single GRPO-style update at both the response and token levels and uses the signed teacher--old-policy token gap to redistribute the verifier-derived component.

Jie Zhang, Jing-Xiao Yang, Zhehao Huang et al. · 1 citation
Preprint Aug 2026

I-SDPO: Instance-Level Adaptive Self-Distillation Policy Optimization

I-SDPO (Instance-Level Adaptive Self-Distillation Policy Optimization), which treats teacher reliance as capability-dependent and uses imitation only where group-relative rewards are uninformative, obtains the best result in all four scientific domains.

Yubo Zhang, Xin-Hong Ma, Zezhong Tan et al. · 1 citation
#artificial intelligence Preprint Sep 2026

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses, improves average performance over FlowRL on both Qwen3-4B and Qwen3-8B, while also improving training speed and stability, avoiding direct OPSD's response-length collapse, and exhibiting higher correc...

Zi-Xun Huang, Kishan Panaganti, Hai-Tao Mi et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning

On-policy self-distillation (OPSD) trains a question-only student with token-level feedback from a teacher given training-only privileged information (PI). OPSD therefore provides dense, on-policy supervision, and is free of a larger external teacher, but its effectiveness rests on how PI is designed and utilized. Our...

Jia-Cheng Du, Wei-Wei Xie, Tian-Yi Du et al. · 0 citations
#machine learning Preprint Sep 2026

Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs

Fire (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning, is proposed, which provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where s...

Seohyun Lee, Dong-Jun Han, Seyyedali Hosseinalipour et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.