Skip to content

On-Policy Distillation for Vision-Language Model Adaptation, an Effective Paradigm on Low-Quality Multimodal Data

Sep 2026 · Information Fusion · Vol 139, pp. 104812 · 0 citations · 44 references
Computer Science

TL;DR

OnPoKD is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision, and is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision.

Abstract

Knowledge distillation offers an efficient route to transfer a task-adapted vision-language teacher to a compact student. The training target in current vision-language distillation methods is typically constructed from the teacher prediction and applied uniformly to all training samples, making it unreliable under class and domain shifts. In this paper, we argue that distillation target construction should be treated as a dynamic training decision rather than a fixed recipe. To this end, we propose OnPoKD, an on-policy distillation framework for vision-language model adaptation. To the best of our knowledge, OnPoKD is the first framework that applies on-policy distillation to vision-language model adaptation by learning target construction as a policy decision. OnPoKD learns a lightweight controller that constructs sample-wise adaptive targets using reliability and disagreement cues from the teacher model, student model, and zero-shot prior. Instead of relying on a fixed teacher prediction, the controller dynamically balances teacher supervision, zero-shot prior guidance, and hard-label anchoring through bounded policy actions, allowing the distillation target to adapt to varying sample reliability and training stages. The policy controller is updated with validation feedback, encouraging target construction to optimize transferability rather than merely fitting the training distribution. Since the controller is only used during training, OnPoKD can be seamlessly integrated into existing vision-language distillation pipelines while preserving the original inference architecture and test-time cost. Extensive experiments on Base-to-novel generalization and Cross-dataset transfer benchmarks show that OnPoKD consistently improves over strong vision-language distillation baselines.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

TTRSD: Test-Time Reinforcement Learning with Self-Distillation for Vision-Language Models

Test-time reinforcement learning enables vision-language models (VLMs) to adapt using unlabeled inputs. However, repeated sampling under fixed visual conditions can reinforce shared perceptual errors, while sequence-level rewards fail to isolate visual perception the foundational bottleneck that anchors multimodal reas...

Shu-Ning Wang, Zhi-Heng Wu, Xun-Lan Zhou et al. · 0 citations
Preprint Sep 2026

CA-OPD: Confidence-Aware On-Policy Distillation for Structured Visual Prediction

Autoregressive vision language models unify heterogeneous perception tasks but are highly susceptible to compounding errors. On-policy distillation (OPD) bridges the training-inference mismatch by training students on their own rollouts. However, unreliable student predictions, especially early in training, can derail...

Meng-Hao Li, Lin-Jie Mu, Yin Wang et al. · 1 citation
#small language model Preprint Aug 2026

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

The role of a frozen off-the-shelf instruct model as the teacher in on-policy distillation is investigated, and a key insight is revealed: the teacher reshapes the student's policy distribution so that subsequent RL converges to a superior solution that RL alone cannot reach.

Qi Ye, Zhi-Yuan Gu, Jingjie Xia et al. · 3 citations
#artificial intelligence Preprint Sep 2026

Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs

Multi-Teacher Self-Distillation Policy Optimization is introduced, an on-policy distillation method that unifies several frozen teachers into one student model and lifts the weakest domain of Qwen3-8B by 14.79 points and narrows its domain gap by 74.7%, a better balance than serving one matched teacher per domain.

Xi-Xiang He, Xing-Ming Li, Bai-Qi Wu et al. · 2 citations
#small language model Preprint Sep 2026

SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show...

Ahmadreza Jeddi, En-Ming Zhang, Jasper Gerigk et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.