This work proposes SkillMaster, a training framework that teaches agents to create new skills, refine existing skills, and select accumulated skills during task solving, and introduces DualAdv-GRPO, which separately estimates advantages for task-solving actions and skill-editing decisions, stabilizing joint training ac...
Min Yang, Jing-Hua Piao, Xuan-Ye Xia et al.· arXiv.org· 7 citations
LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowle...
Min Yang, Yi-Chen Pan, Jing-Hua Piao et al.· 0 citations
Outcome-Verified Comparative Self-Distillation (OVCSD) is proposed, which organizes failed student rollouts into a prefix tree, adaptively invokes a skill-conditioned teacher from student-reached states, and retains only outcome-verified successful continuations.
Xuanye Xia, Jinghua Piao, Min Yang et al.· arXiv.org· 0 citations
The results suggest that the sweet spot for small models in large-model inference systems lies not in solving complex tasks independently, but in performing lightweight, structured, and verifiable auxiliary operations.
Jingquan Chen, Jie Feng, J. Piao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.