Sep 2026· Information Fusion· Vol 139, pp. 104812· 0 citations· 44 references
Computer Science
TL;DR
OnPoKD is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision, and is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision.
Abstract
Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in current vision-language distillation methods is typically constructed from the teacher prediction and applied uniformly to all training samples, making it unreliable under class and domain shifts. In this paper, we argue that distillation target construction should be treated as a dynamic training decision rather than a fixed recipe. To this end, we propose OnPoKD, an on-policy distillation framework for vision-language model adaptation. To the best of our knowledge, OnPoKD is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision. OnPoKD learns a lightweight controller that constructs sample-wise adaptive targets using reliability and disagreement cues from the teacher model, student model, and zero-shot prior. Instead of relying on a fixed teacher prediction, the controller dynamically balances teacher supervision, zero-shot prior guidance, and hard-label anchoring through bounded policy actions, allowing the distillation target to adapt to varying sample reliability and training stages. The policy controller is updated with validation feedback, encouraging target construction to optimize transferability rather than merely fitting the training distribution. Since the controller is only used during training, OnPoKD can be seamlessly integrated into existing vision-language distillation pipelines while preserving the original inference architecture and test-time cost. Extensive experiments on Base-to-novel generalization and Cross-dataset transfer benchmarks show that OnPoKD consistently improves over strong vision-language distillation baselines.
Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reas...
Shu-Ning Wang, Zhi-Heng Wu, Xun-Lan Zhou et al.· 0 citations
Autoregressive vision language models unify heterogeneous perception tasks but are highly susceptible to compounding errors. On-policy distillation (OPD) bridges the training-inference mismatch by training students on their own rollouts. However, unreliable student predictions, especially early in training, can derail...
Meng-Hao Li, Lin-Jie Mu, Yin Wang et al.· 1 citation
The role of a frozen off-the-shelf instruct model as the teacher in on-policy distillation is investigated, and a key insight is revealed: the teacher reshapes the student's policy distribution so that subsequent RL converges to a superior solution that RL alone cannot reach.
Qi Ye, Zhi-Yuan Gu, Jingjie Xia et al.· 3 citations
Multi-Teacher Self-Distillation Policy Optimization is introduced, an on-policy distillation method that unifies several frozen teachers into one student model and lifts the weakest domain of Qwen3-8B by 14.79 points and narrows its domain gap by 74.7%, a better balance than serving one matched teacher per domain.
Xi-Xiang He, Xing-Ming Li, Bai-Qi Wu et al.· 2 citations
LT-OPD, a training framework for extreme visual-token reduction, is proposed and it is shown that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.
Junxian Li, Rui-Xuan Yang, Tian-Ao Zhang et al.· 0 citations
Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show...
Ahmadreza Jeddi, En-Ming Zhang, Jasper Gerigk et al.· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduSep 24, 2026
A new method, called CW-Net, translates the reasoning process of an autonomous vehicle’s AI system into understandable concepts that explain its behavior.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.