Recursive self-improvement (RSI) allows agents to carry useful changes across generations. Maintaining safety across these generations involves both preventing unsafe behavior from persisting and enabling recovery when failures occur. We study these challenges through a controlled testbed of stateful authorization task...
Yun-Bei Zhang, Janet Wang, Saiyue Lyu et al.· 0 citations
Harness evolution improves an LLM agent's prompts, tools, and workflow, while the optimizer's own tools and procedures often remain fixed. We study whether an optimizer can improve another agent more effectively by also improving how it diagnoses failures, develops edits, and tests their effects. Two observations guide...
Ze-Kai Wang, Ying-Qiang Ge, Ze-Kun Wang et al.· 0 citations
Multi-agent systems derive their capabilities from sharing evidence, delegating tasks, and combining information across agents. The same process creates a safety problem: contributions that are admissible in isolation can jointly enable a prohibited use. Blocking every sensitive action avoids disclosure but defeats the...
Yun-Bei Zhang, Saiyue Lyu, Janet Wang et al.· 0 citations
The main idea is to use all the negative samples when optimizing the learning objective to avoid the sampling process, and rearrange the origin loss function into a linear form and take advantage of meticulous mathematical derivation to reduce the complexity of the loss function.
Pierpaolo Basile, Birgitta Dresp-Langley, Jianchao Ji et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.