Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Aug 2026

A Critical Review of LLM Agents for Automated Penetration Testing: Benchmark Realism and Evidence Grading

Penetration testing has long resisted full automation because it requires contextual reasoning, adaptive tool use, and experience-driven decision making across the Penetration Testing Execution Standard (PTES) lifecycle. Recent advances in large language models (LLMs), agentic reasoning, and multi-tool invocation have shifted the field from rule-based scripts and attack-graph reinforcement-learning planners toward interactive agents capable of planning, reflection, and environment interaction. The accompanying literature, however, suffers from inconsistent evaluation protocols, opaque scaffolding, and indiscriminate mixing of peer-reviewed papers with vendor self-reports, rendering headline claims difficult to compare. This work presents a critical narrative review with structured evidence charting of AI-assisted automated penetration testing between January 2023 and May 2026, retaining DeepExploit and AutoPentest-DRL as historical anchors. We introduce a dual-axis framework that annotates every finding by environment realism (R1–R4, from single-challenge CTFs to live enterprise networks) and source authority (A–D, from peer review to vendor self-report). Frontier agents reach 22–44% on R1 single-challenge CTFs (D-CIPHER), 79.17% on R2 AutoPenBench subtasks with a domain-tuned 32B model (xOffense), and 71.4% on the AISI 95-task expert tier (GPT-5.5), yet only 20–30% end-to-end completion on the R3 32-step TLO cyber range. On R4 live enterprise networks, the ARTEMIS multi-agent framework placed second overall against ten human professionals on an 8,000-host university network, submitting nine valid vulnerabilities at an 82% acceptance rate. Robust autonomy remains weak in long-horizon exploitation, Active Directory lateral movement, and defender-present environments. We argue that automated penetration testing is best framed as a systems-engineering problem rather than a single-model race, and we distinguish engineering-tractable Type A capability gaps from planning-bottlenecked Type B failures that require base-model reasoning improvements or reinforcement-learning post-training.

Yinhang Zhou, Weihua Xu · 0 citations