While LLM-based attackers exhibit growing proficiency in vulnerability exploitation, most existing cybersecurity benchmarks suffer from single-stage truncation, prematurely terminating evaluation upon initial access. In practice, initial footholds are exceptionally fragile across operational disruptions such as service...
Su-Jin Chen, Lijun Li, Xu-Hong Wang et al.· 0 citations
MisKnow-Agent is introduced, a controlled evaluation framework that constructs task-specific documents supporting manually audited false conclusions with controlled authority cues and source styles that evaluates DeerFlow and WebThinker with three backbone LLMs using a report-level false-conclusion adoption rate that c...
Peng-Yu Zhu, Lijun Li, Long-Ping Yang et al.· arXiv.org· 4 citations
This work presents UniACE, a unified framework for model-centric evaluation under an explicit, common execution condition, and reports agent benchmark outcomes as properties of an explicit evaluation configuration, enabling more interpretable and reproducible cross-benchmark comparisons.
Peng-Yu Zhu, Lijun Li, Yaxing Lyu et al.· 3 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.