Harness evolution improves an LLM agent's prompts, tools, and workflow, while the optimizer's own tools and procedures often remain fixed. We study whether an optimizer can improve another agent more effectively by also improving how it diagnoses failures, develops edits, and tests their effects. Two observations guide...
Ze-Kai Wang, Ying-Qiang Ge, Ze-Kun Wang et al.· 0 citations
The model's hidden state provides a way to read, steer, and check tool choice before a call is made, suggesting that the model's hidden state provides a way to read, steer, and check tool choice before a call is made.
Ze-Kun Wu, Ze-Kun Wang, Seonglae Cho et al.· arXiv.org· 11 citations· ⚡2
A three-tier diagnostic taxonomy shows that verification/feedback and planning failures dominate execution/grounding errors, while a single scalar success rate can not explain, and connects these findings to newer long-horizon CUA benchmarks and derive stage-specific design rules for CUA evaluation.
Zi-Han Dong, Zhiyuan Ma, Zekun Wang et al.· arXiv.org· 2 citations
The architecture and the Muon optimizer together shift the optimal learning rate and batch size upwards, render batch-size warmup unnecessary, and substantially improve stability under stress tests.
Zi-Han Qiu, Ze-Kun Wang, Xiao Li et al.· 24 citations· ⚡3
Qwen-CUA is introduced, a native computer-use agent with a 397B-A17B Qwen mixture-of-experts backbone that outperforms Qwen3.7 and remains competitive with leading proprietary systems, and scalable verifiable interaction and hybrid tool use as key directions.
Dunjie Lu, Shuai Bai, Tianyi Bai et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.