This work proposes SkillMaster, a training framework that teaches agents to create new skills, refine existing skills, and select accumulated skills during task solving, and introduces DualAdv-GRPO, which separately estimates advantages for task-solving actions and skill-editing decisions, stabilizing joint training ac...
Min Yang, Jing-Hua Piao, Xuan-Ye Xia et al.· arXiv.org· 7 citations
GARDiff, a Graph-Aligned Residual Diffusion framework for probabilistic multivariate time-series forecasting,Instead of treating deterministic-derived graphs as fixed diffusion conditions, GARDiff progressively adapts them to residual generation, improving probabilistic forecasting performance and uncertainty calibrati...
Rui-Na Han, Min Yang, Xu Zhang et al.· 0 citations
LLM agents are increasingly expected to support enterprise workflows, where tasks often involve missing information, uncertainty, feedback, and long-term trade-offs. However, existing enterprise and financial benchmarks mainly test static capabilities such as information extraction, numerical calculation, domain knowle...
Min Yang, Yi-Chen Pan, Jing-Hua Piao et al.· 0 citations
Safety alignment trains large language models to refuse harmful requests stated plainly, but that training is applied mostly to surface form. Requests that only recontextualise the same operational content, changing how the model reads it, are therefore only weakly covered. The ASCII Attack is one such recontextualisat...
Dayong Gu, Yi-Fei Dong, Xing-Hao Yang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.