This work proposes SkillMaster, a training framework that teaches agents to create new skills, refine existing skills, and select accumulated skills during task solving, and introduces DualAdv-GRPO, which separately estimates advantages for task-solving actions and skill-editing decisions, stabilizing joint training ac...
Min Yang, Jing-Hua Piao, Xuan-Ye Xia et al.· arXiv.org· 7 citations
Outcome-Verified Comparative Self-Distillation (OVCSD) is proposed, which organizes failed student rollouts into a prefix tree, adaptively invokes a skill-conditioned teacher from student-reached states, and retains only outcome-verified successful continuations.
Xuanye Xia, Jinghua Piao, Min Yang et al.· arXiv.org· 0 citations
This work presents a stable and efficient training pipeline, incorporating algorithmic and system optimizations such as clipped importance sampling, training-inference ratio correction, and mixed-precision control, and proposes a structured evaluation framework across three dimensions: comprehensibility, reproducibilit...
Xinyu Tang, Gangqiang Cao, Yurou Liu et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.