Agent Skills, the SKILL.md files that tell an LLM coding agent how a project works, are revised like code, yet what a revision does to the agent is unknown. From 2,608 first/last revision pairs of 3,159 Skills, we characterize how Skills evolve and how they change together with the configuration of the agent's harness....
Jia-Jie Wang, Yu-Tong Zhao, Tian-Lin Li et al.· 0 citations
KGMACG is evaluated on three industrial-scale case studies against six state-of-the-art multi-agent baselines: MetaGPT, AutoGen, CAMEL, CrewAI, ChatDev and CodeAgent and indicates that KGMACG advances the automation of application-level software development.
Bo Yang, Xiao Zhang, Weisong Sun et al.· ACM Transactions on Software...· 0 citations
A format-aware metamorphic testing framework with three metamorphic relations is proposed to comprehensively evaluate the format robustness of end-to-end LLM document workflows and demonstrates that document format is not a neutral wrapper but a critical factor affecting the reliability of LLM software systems.
Xiaoyu Zhang, Xianyun Cheng, Tianlin Li et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.