Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, lea...
P. Wang, Chen-Hao Liang, Ze-Long Xu et al.· 0 citations
WeClawArena is introduced, an auditable benchmark and runtime sandbox for multi-party owned-agent collaboration over personal workspaces and audits attack success from bounded runtime evidence, supporting diagnosis of task breakdown, privacy leakage, poisoned evidence, and invalid authority paths.
P. Wang, Ao-Jie Yuan, Haiyu Zhang et al.· 1 citation
None of the step-level credit signals the authors audit -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- shows reliable incremental fidelity beyond its own marginal-matched shuffled control.
Haiyu Zhang· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.