On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories. Although OPD performs strongly when teacher and student belong to the same model family, we find that its effectiveness degrades in cross-family settings even after tokenizer alignment, with substantially stronger ext...
Nai-Bin Gu, Qing-Yi Si, Chen-Xu Yang et al.· 0 citations
Focused On-demand Visual Evidence Adaptation is proposed, a cache-friendly approach that builds a reusable visual memory and dynamically retrieves a bounded subset for a draft state and demonstrates that state-conditioned evidence retrieval is an effective alternative to reusing a fixed visual representation throughout...
He Zhu, Da-Yan Wu, Zihao Zhang et al.· 0 citations
CoRe-MoE is proposed, a Compact Reusable MoE framework for parameter-efficient continual multimodal instruction tuning that improves final average performance over the strongest competing baseline by up to 5.90 points, while using less than 1% of the trainable parameters required by sequential LoRA for later tasks.
Run-Ze Liu, Naibin Gu, Ming-Xu Ai et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.