In harness self-evolution, agents modify their own prompts, code, tools, and orchestration while keeping the underlying language model fixed. Recent work has shown that agents can improve themselves in response to task failures and achieve substantial performance gains. However, gains on failed tasks do not automatical...
Qi Cai, Yong-Gang Zhang, Jun Nie et al.· 2 citations· ⚡2
This work constructs a specialized dataset to demystify the emotional circuits underlying the three-stage ``Adapt-Aggregate-Execute''mechanism and discovers a functional decoupling: visual emotional cues are aggregated in middle layers via sentiment-specific attention heads, but are subsequently translated into narrati...
EvolveNet is introduced, a paradigm of collaborative harness evolution that moves experience extraction to the data and introduces scope-typed, evidence-guided program aggregation, which improves the shared harness in all five settings.
Jun Nie, Yong-Gang Zhang, Qi Cai et al.· 2 citations
DRNOISE, a 100-task benchmark for answer recovery under misleading evidence, is introduced, a 100-task benchmark for answer recovery under misleading evidence that requires active reconciliation of direct claims with record-level evidence.
Jun Nie, Zhiqin Yang, Zhenheng Tang et al.· arXiv.org· 2 citations
This work establishes a new paradigm for generated image detection by recasting the detection task as a problem of machine unlearning, and introduces two detection methods: data-free detection, which prunes model parameters to induce unlearning without data access, and data-driven detection, which optimizes LVMs to unl...
Jun Nie, Yonggang Zhang, Tongliang Liu et al.· 0 citations
It is proved that the planner's suboptimality is bounded by twice this discrepancy between the predicted and the true plan-cost at the plan the planner commits to, whereas the data-averaged prediction error neither bounds nor tracks it.
Hanzhe You, Yonggang Zhang, Maohao Ran et al.· arXiv.org· 2 citations
Test-Time Harness Evolution is introduced, which treats the executable harness as the state of test-time adaptation for LLM agents as evolution over executable control programs and identifies execution-derived proxy reliability as a central challenge for robust unsupervised agent improvement.
Jun Nie, Yonggang Zhang, Jun Song et al.· arXiv.org· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.