Skip to content

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

CyberPersistBench: Evaluating LLM-Based Cyber Attackers on Installation and Persistence

While LLM-based attackers exhibit growing proficiency in vulnerability exploitation, most existing cybersecurity benchmarks suffer from single-stage truncation, prematurely terminating evaluation upon initial access. In practice, initial footholds are exceptionally fragile across operational disruptions such as service...

Su-Jin Chen, Lijun Li, Xu-Hong Wang et al. · 0 citations
Jul 2026

Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions

MisKnow-Agent is introduced, a controlled evaluation framework that constructs task-specific documents supporting manually audited false conclusions with controlled authority cues and source styles that evaluates DeerFlow and WebThinker with three backbone LLMs using a report-level false-conclusion adoption rate that c...

Peng-Yu Zhu, Lijun Li, Long-Ping Yang et al. · 4 citations
#artificial intelligence Preprint May 2026

UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

This work presents UniACE, a unified framework for model-centric evaluation under an explicit, common execution condition, and reports agent benchmark outcomes as properties of an explicit evaluation configuration, enabling more interpretable and reproducible cross-benchmark comparisons.

Peng-Yu Zhu, Lijun Li, Yaxing Lyu et al. · 3 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.