Self-Evolutionary Reinforced Knowledge Distillation for Multi-Modal Tool-Use Agents
A two-stage self-evolutionary knowledge distillation framework that equips small MLLMs with robust and adaptive tool-use behaviors and introduces weighted semantic objectives and iteratively expand competence through error-driven optimization, hybrid experience replay, and group-relative policy refinement with multi-di...